DDC Assurance Lab is not an accredited laboratory.

AI & agentic systems

Assurance for systems that can act.

Agentic systems combine model behavior with tools, credentials, state, external services and execution authority. DDCAL evaluates that wider technical boundary rather than treating model output alone as the system.

AI agent assurance beyond model evaluation

Traditional model evaluation can measure response quality, benchmark performance or policy compliance. An operational AI agent introduces additional questions: which actions may it take, what tools can it invoke, what resources can it affect, what state does it inherit, and how is successful execution independently verified?

DDC Assurance Lab structures technical assessment around those boundaries. The objective is not to claim that an agent is universally “safe.” The objective is to test defined claims under an explicit scope and preserve the evidence, limitations and failure conditions needed to understand the result.

What DDCAL can examine

  • Delegated authority: whether requested outcomes, retrieved data or model output can silently expand permission.
  • Tool and privilege boundaries: whether capabilities are constrained to the authority actually granted.
  • State transitions: whether effectful actions are bound to the correct predecessor and produce a verifiable successor state.
  • Verifier independence: whether a producer, agent or shared dependency is effectively validating its own result.
  • Evidence and provenance: whether important claims remain traceable to source, execution and verification records.
  • Recovery behavior: whether rollback and recovery can be accepted without trusting a component already under suspicion.

Typical assessment questions

Examples include whether an agent can invoke a high-impact tool without an immutable intent binding, whether retries can accidentally duplicate an effect, whether approval state can be confused with correlation state, whether generated evidence can be trusted independently of the producing component, and whether a failure can be recovered without silently promoting stale or self-certified state.

Result boundaries matter

A DDCAL result applies only to the stated implementation, version, environment, requirements and evidence boundary. A PASS for one control does not automatically establish system-wide safety, regulatory compliance or future behavior. Where evidence is insufficient, the appropriate result is INCONCLUSIVE or NOT ASSESSED rather than an expanded claim.

Discuss an agent assurance assessment