Same approval. Missing evidence link. Configured CI gate blocked.
agent-assure catches regressions in declared, structured process evidence that
final-answer-only checks can miss. It turns privacy-filtered run records into
reproducible comparisons, reviewer-facing artifacts, portable evidence packets,
and ordinary CI gate signals. Evidence strength follows its recorded origin:
fixture values are authored test inputs, instrumented-adapter values remain
producer-attested, and direct model responses do not independently prove that a
tool, evidence lookup, policy check, or human review occurred.
Local-first · offline flagship demo · versioned artifacts · CI-native · no hosted control plane required
Run the offline demo · Inspect the reviewer output
For AI leaders · For architects · For engineers
10 deterministic fixture cases · 0 decision fields changed ·
claim-duration: linked → missing · classification: new_failure ·
configured CI gate: blocked
Note
This is a bundled deterministic fixture demonstration, not a model benchmark, live-model result, or customer outcome. It requires no provider API key, network call, or token spend.
Read the flagship demonstration
| Your role | Start with the question that matters |
|---|---|
| Chief AI Officers and release owners | Did a release preserve its declared controls? See the local evidence and portable handoff for repeatable review. Read the AI leader brief. |
| AI/ML architects and platform owners | Where does it fit, what crosses the trust boundary, and which contracts are stable? Review the architecture · Inspect the API surface. |
| AI/ML engineers | How do I turn observable expectations and privacy-filtered run records into a configured CI decision? Follow the engineering guide. |
A final-answer-only check can see no decision-field change while a declared release expectation regresses:
- Evidence and RAG support: a required source or material claim-to-evidence
link disappears, or declared corpus and retrieval identity changes. Example:
MATERIAL_CLAIM_MISSING_EVIDENCE → new_failure. The development surface also includes an authority-scoped, deterministic controlled evidence-sensitivity contract. Its v1 producer is a declarative synthetic fixture harness, not evidence that a model used contextual evidence instead of parametric memory. A separate repeated paired protocol supports the same narrow binary response relation for prebound stochastic live studies. Its fixed planned-frame composite rate includes every frozen cluster, pessimistically scores a non-analyzable cluster as zero, and reports observed/analyzable counts separately. Verdict-bearing packets and graphs bind the exact candidate RunSet, while both source execution configurations remain anchored to their predeclared arms. Conclusions remain conditional on statistical sufficiency, exact subject binding, and declared cluster assumptions. The untagged development surface also provides a preregistered real-model study workflow, but this repository contains no real-provider study result. - Human review: a required route or performed-review record is missing.
- Provider, tool, and privacy boundaries: a forbidden provider or tool appears, or declared route, redaction state, or detector identity changes.
- Usage and reliability: retries, tool calls, tokens, latency, rate-limit events, or declared estimated cost change materially.
- Streaming integrity: events are replayed, duplicated, conflicting, or outside the declared sequence contract.
- Protocol-bound live behavior: repeated observations drift outside a declared protocol or comparison boundary, or a paired evidence-response study is underpowered, structurally invalid, or inconsistent with its exact pre-execution arm bindings.
A surfaced difference may be blocking, review-only, or informational. Usage and reliability deltas block only when a suite or policy declares that behavior.
Requires Python 3.11 or newer.
pip install agent-assure
agent-assure demo flagship --out .tmp/demo/flagship --cleanThe installed package runs the bundled deterministic fixture from any directory—no repository clone, provider API key, network call, or token spend.
output equivalence: preserved
missing evidence link: claim-duration
reason code: MATERIAL_CLAIM_MISSING_EVIDENCE
classification: new_failure
CI gate: blocked as expected
The demo wrapper exits 0 only when it verifies that the expected regression
was caught. The underlying candidate evaluation, comparison, and CI commands
remain strict and exit nonzero for the blocking finding.
Screenshot of the reviewer-facing
evidence-diff.html produced from the same bundled fixture. Open the
image to inspect it at full resolution.
Key artifacts are written under .tmp/demo/flagship:
| Artifact | Review purpose |
|---|---|
demo-summary.json |
Machine-readable demonstration result |
baseline-report/evaluation-summary.json |
Baseline behavior against declared expectations |
comparison-report/comparison-summary.json |
Controlled baseline-to-candidate classification |
ci-report/evidence-packet.json |
Portable machine-readable review handoff |
evidence-diff.html |
Self-contained human-readable evidence diff |
How this README evidence view is verified against the fixtures
This diagram is checked in CI against the bundled flagship fixtures, keeping README claims aligned with the evidence produced by the project itself.
flowchart LR
subgraph OutputCheck["Ordinary visible-output check"]
BOut["Baseline output<br/>recommendation=approve<br/>outcome=approve"]
COut["Candidate output<br/>recommendation=approve<br/>outcome=approve"]
Same["Visible answer unchanged"]
BOut --> Same
COut --> Same
end
subgraph InvariantCheck["agent-assure invariant check"]
BEv["Baseline evidence<br/>claim-duration linked"]
CEv["Candidate evidence<br/>claim-duration missing link"]
Pass["Baseline evaluation: pass"]
Fail["Candidate evaluation: fail<br/>MATERIAL_CLAIM_MISSING_EVIDENCE"]
BEv --> Pass
CEv --> Fail
end
Same --> Tension["Output unchanged<br/>but governance invariant regressed"]
Equiv["Fixture equivalence: pass"] --> Compare["Baseline-to-candidate comparison"]
Pass --> Compare
Fail --> Compare
Tension --> Compare
Compare --> NewFailure["Classification: new_failure"]
agent-assure complements the evaluation, observability, runtime-control, and
governance systems teams already use.
| Layer | Primary question | Relationship to agent-assure |
|---|---|---|
| Output and agent evals | Does the answer, trajectory, tool use, or component meet its quality criteria? | Adds source-aware checks for declared process-evidence expectations. |
| Observability and tracing | What happened during execution? | Consumes versioned, privacy-filtered evidence; it is not a telemetry backend. |
| Runtime guardrails | What must change or stop during a request? | Evaluates at release time; it is not runtime enforcement. |
| Governance and GRC systems | Which policies, approvals, and accountabilities apply? | Supplies review evidence; it is not a system of record and does not determine compliance. |
agent-assure |
Did a controlled candidate preserve declared process expectations? | Evaluates, compares when equivalent, packetizes, and returns a CI signal. |
It is a particularly strong fit when release review must be local, reproducible, CI-enforceable, and traceable without a required hosted control plane.
Declare → Capture from a declared source (privacy-filtered) → Evaluate
→ Compare (when equivalent) → Packet → Gate
Declared expectations and canonical run evidence remain distinct. The candidate is evaluated first; equivalent-baseline context is added only after comparison prerequisites pass. The evidence packet then supports a CI signal and human release review.
agent-assure is a local-first Agent Release Assurance Compiler: controls and
evidence stay bound to method identity, prerequisites, provenance, assumptions,
and limitations.
The assurance model is deliberately bounded:
- Deterministic and reproducible in fixture mode: fixed, versioned fixtures, canonical serialization, schemas, and digest-bound manifests make checks repeatable; that reproducibility does not estimate production prevalence.
- Scoped invariance claims: results cover only declared, observable fields and explicit prerequisites—not hidden reasoning or all production behavior.
- Traceable lineage: expectations connect to RunSets, findings, comparisons, evidence packets, and configured gate state; provenance records participating material.
- Fail-closed: malformed, conflicting, incompatible, ambiguous, or unbound evidence does not silently become a passing review.
- Statistically bounded: live conclusions about probabilistic provider behavior remain tied to a declared statistical protocol, with its data boundary, configuration, window, sampling noise, dependence, and limitations explicit.
Architecture choices and evidence boundaries are documented in architectural decision records (ADRs), including deterministic fixture versus stochastic live semantics.
The development-RFC core/v1 mutation catalog runs seven deterministic
challenges across evidence linkage, human-review routing, tool boundaries,
provenance identity, privacy redaction, duplicate replay, and budget-stop
integrity:
agent-assure controls mutate \
--suite assurance/suite.yaml \
--runset runs/baseline.json \
--catalog core/v1 \
--seed 0 \
--today 2026-08-02 \
--full-report \
--out reports/control-challengeEvery selected operator runs independently against the same immutable source.
The output binds the canonical catalog digest, normative expected detector,
observed and prohibited substitute findings, exact changed paths, provenance,
independence class, seed, and limitations. The campaign itself remains a
finite challenge record. controls efficacy can derive an exact
catalog-relative detector kill ratio over completed applicable outcomes while
preserving inapplicable, invalid, and error counts outside its denominator.
That ratio is not a safety score, statistical confidence interval, or
universal-coverage claim.
Only deterministic caught and survived campaign outcomes may contribute a
control-efficacy verdict. An applicable critical threat with no completed
challenge emits CRITICAL_THREAT_UNCOVERED and requires review under the
default profile. Required and critical survivors, invalid/error outcomes, and
required non-verdict outcomes are always blocking. Evidence-packet ci gate
requires control-efficacy evidence by default and uses strict verification when
it is present; pass a verifier-owned controls-mutation YAML with
--efficacy-policy. --require-efficacy remains as an explicit restatement
for existing automation. Strict CI accepts only complete, all-caught,
all-applicable-challenged evidence and pins the catalog, selected and required
operators, and threat manifest. The conspicuous
--allow-missing-efficacy-for-migration escape hatch is only for temporary
non-assurance migration and records that weaker profile. Use
--allow-advisory-efficacy only for an explicit review flow with evidence
present.
For an efficacy-bearing release claim, ci gate --release-profile additionally
requires an evidence packet, a verifier-owned --efficacy-policy, present
efficacy evidence, and blocking warning/not-evaluated handling. It rejects the
advisory and compatibility weakening flags. This is a CI efficacy profile, not
publication authorization. make release-publish-check runs that strict
profile against the separately staged release efficacy packet and verifier
policy before the empirical-readiness and engineering release checks; all are
necessary, and none alone authorizes publication.
The packaged offline demonstration exercises both a strong and deliberately weakened assurance control, then verifies that an unrelated failure cannot substitute for the expected detector:
agent-assure demo assure-the-assurance \
--out .tmp/demo/assure-the-assurance \
--cleanFor a minimal editable workflow, start with agent-assure init controls-mutation, run the read-only doctor controls-mutate preflight, then
produce a campaign and control-efficacy-report.json.
Inspect the exact seven-operator catalog · Measure control efficacy · Read "Who assures the assurance?" · Review the evidence contracts
agent-assure integrates through declared YAML expectations, versioned run
evidence, the documented CLI, and the framework-neutral AgentRunRecord
producer contract.
| You provide | agent-assure does |
You receive |
|---|---|---|
| Declared expectations, a candidate RunSet, and an optional equivalent baseline RunSet | Validate, evaluate, compare when equivalence prerequisites pass, packetize, and apply the configured gate | Evaluation and comparison summaries, evidence-packet.json, human-readable reports, and an ordinary CI exit status |
The integration contract has three parts:
- Declare structured process-evidence expectations in YAML.
- Produce versioned run records from fixtures or privacy-filtered adapters with explicit field origins.
- Evaluate the candidate, compare equivalent baseline evidence when available, and gate the resulting evidence packet.
A real expectation from the bundled flagship suite:
cases:
- case_id: shared-source-multi-claim
fixture_id: shared-source-multi-claim
expectation:
expected_recommendation: approve
required_evidence_refs:
- ref-shared-clinical-note
material_claim_ids:
- claim-eligibility
- claim-durationagent-assure does not infer material claims from rationale text. Authors
declare the oracle, and run-record producers emit explicit claim-to-evidence
links for the material claims they intend to satisfy.
Author expectations · Understand the CLI contract · Review the public API surface · Understand evidence packets
GitHub Actions example using the bundled fixture
Pin both the package and composite action in release workflows. The example uses the latest published tag, v0.6.5; move both pins together only after a newer tag is published. Replace the example suite and variant paths with your own controlled materials.
name: agent-assure
on: [pull_request]
jobs:
assure:
runs-on: ubuntu-latest
steps:
- uses: actions/checkout@v4
- uses: actions/setup-python@v5
with:
python-version: "3.11"
- run: python -m pip install agent-assure==0.6.5
- uses: acblabs/agent-assure/.github/actions/agent-assure@v0.6.5
with:
suite: examples/prior_auth_synthetic/suite.yaml
baseline-variant: examples/prior_auth_synthetic/variants/baseline.yaml
candidate-variant: examples/prior_auth_synthetic/variants/candidate_evidence_normalization.yaml
report-mode: fullfull produces the complete review artifacts; fail-fast gives shorter
blocking feedback. The published v0.6.5 action shown above predates the
efficacy-required default and is a fixture smoke example, not evidence of
control efficacy. After v0.6.6 or a later version is published, move both pins
together. Its composite action fails closed because it does not construct
control-efficacy evidence; fixture-only, evaluation-only migration jobs must
explicitly set allow-missing-efficacy-for-migration: "true". An assurance
workflow must instead construct an efficacy-bearing packet and gate it with a
separate verifier-owned policy. The configured gate follows declared
expectations and policies, the selected gate profile, and explicit strictness
flags.
The composite action uploads only the packet, its privacy-filtered assurance
evidence graph, manifest, summaries, and CI diagnostics by default. Set
upload-full-artifacts: "true" only when the workflow is approved to retain
compiled suites, fixture data, and RunSets; the default retention period is 14
days.
Current published release: v0.6.5 on GitHub and PyPI. This checkout is the
unreleased 0.6.6 candidate and must not be described or installed as a
published release until the empirical publish gate passes.
The CLI, YAML authoring format, persisted versioned JSON artifacts, and
AgentRunRecord producer contract are the primary integration surface.
Framework adapters, streaming, and live execution remain experimental.
The RC label applies only to the primary surface; development-RFC contracts
remain non-stable. PyPI's Development Status :: 4 - Beta is the closest
standardized classifier to an RC and does not widen that surface.
| If you have… | Start with… | Maturity |
|---|---|---|
| YAML suites or versioned JSON artifacts | CLI contract | Primary supported surface |
| A GitHub release workflow | Composite action | Packaged and documented |
| Deterministic mutation operators and closed catalog campaigns | Core mutation catalog · Evidence-carrying releases | Development RFC |
| RAG retrieval evidence | RAG provenance demo | Reference implementation |
| JSONL or multi-agent events | Streaming example | Experimental |
| LangGraph or Google ADK events | LangGraph · Google ADK | Experimental |
| Live provider or external-script subjects | Adapter contract | Experimental, time-bound evidence |
| Preregistered real-model measurement | Real-model study | Untagged development contract; no study result published |
| Independently controlled CI learning pilot | External pilot evidence | Untagged development contract; no qualifying pilot recorded |
| OpenTelemetry context or export | OpenTelemetry alignment | Optional alignment only |
Want to help with the still-unmet external evidence checkpoint? The short external pilot quickstart targets 10–15 minutes of participant effort in a non-maintainer-controlled fork; CI runtime may be longer. It needs no participant-supplied model key, proprietary data, or repository secret. Interest or a workflow run is not evidence completion: capture, participant friction/finalization and prospective consent, and a byte-bound human review are all required for the pilot bundle.
Experimental streaming semantics
Here, idempotency refers only to idempotent deduplication for stable at-least-once redeliveries. Conflicting duplicates fail closed; deterministic sorting prevents out-of-order arrival jitter from changing the persisted trajectory.
Framework adapters project only privacy-filtered agent_assure metadata into
the framework-neutral run-record model. They ignore raw prompts, messages,
completions, tool arguments, token chunks, and unredacted summaries.
Packet-resident evidence can be mapped to selected concepts in the NIST AI RMF, OWASP Top 10 for LLM Applications 2025, ISO/IEC 42001, and MITRE ATLAS 2026.06.
These crosswalks are planning and review aids. They do not establish framework conformance, complete coverage, third-party assurance, or endorsement.
Measured evidence, not a blanket trust claim.
This project is not a compliance attestation.
Generated artifacts make declared inputs, findings, limitations, and gate state
traceable and auditable for human review. Whether the release decision is
defensible remains a human and organizational judgment. agent-assure does not
determine safety or replace domain, legal, regulatory, clinical, security,
provider-quality, model-quality, or business-impact review.
agent-assure is |
agent-assure is not |
|---|---|
| Release-review evidence for declared process expectations | A legal or regulatory determination |
| A deterministic and protocol-bound measurement toolkit | A safety determination |
| A way to surface evidence, routing, privacy, boundary, provenance, usage, and stream-integrity regressions | A general model-quality benchmark |
| A local artifact and CI-gate workflow | A production observability backend or enterprise governance system of record |
| An engineering evidence source for human and governance review | A replacement for organizational accountability |
Pattern redaction is a guardrail, not comprehensive DLP or de-identification. Live conclusions remain bounded by the declared protocol, data boundary, provider/model configuration, and execution window. Review the claim boundary, limitations, threat model, privacy model, and security guidance.
-
Statistical methods: Repeated paired evidence sensitivity · Preregistered real-model study · Live calibration
-
Start: Documentation · For AI leaders · For architects · For engineers
-
Demos: Assure the assurance · Flagship · RAG provenance · Expense approval
-
Integrations: LangGraph · Google ADK · Adapter contract
-
Assurance: What this measures · Control efficacy · Evidence packets · Live calibration
-
Evidence-carrying releases: Core mutation catalog · Minimal evidence graph · Contracts and campaign guide · Architecture · CLI contract
-
Security and governance: Claim boundary · Threat model · Governance crosswalks
-
Project: Contributing · Changelog · License
Development from a repository checkout
pip install -e ".[dev]"
git config core.hooksPath .githooks
python scripts/check_docs_alignment.py
ruff check .
mypy src scripts
pytest
python -m buildThis project ships a CITATION.cff. Use GitHub’s
Cite this repository control for generated citation formats.