Skip to content

Repository files navigation

Surety

A research program for testing whether AI-agent control systems enforce the rules they claim to enforce.

The name is borrowed from nuclear surety: the discipline concerned not simply with whether a system is secure in general, but with whether an unauthorized act is possible at all.

Long-term research question

What evidence justifies trust in an AI-assisted decision system beyond identity and authorization?

Surety is the research repository. The remaining repositories exist to test, demonstrate, or falsify specific aspects of the hypotheses developed here.

The first Surety program tests a narrower part of that question: whether the control plane around an AI agent continues to enforce its stated policy under adversarial pressure. Future work may extend the evidence model to explainable risk, provenance, and hardware-backed assurance without changing the central distinction between authorization and justified trust.

Status: design-phase research repository

Surety currently contains a research question, provisional threat model, benchmark-design constraints, and roadmap. It does not yet contain a released harness, conformance suite, dataset, scorecard, certification program, or validated benchmark result.

AI-agent security is usually evaluated at the model boundary. Can a prompt override instructions? Can a model be induced to misuse a tool? Does a safeguard detect a known attack? Those questions matter, but they leave a second system largely untested: the harness that gives an agent identity, credentials, tools, memory, delegated authority, approval paths, and the ability to act.

Surety asks a narrower question:

When an agent operates under adversarial pressure, does the surrounding control system continue to enforce its stated policy?

The initial goal is an open benchmark specification for agent authorization enforcement. It will define testable invariants, adversarial test families, evidence requirements, and repeatable scoring for the control plane around an agent. It is not a model leaderboard, a product, or a claim that the problem has already been solved.

The problem

An agent can remain inside its nominal permission set and still produce an unauthorized result. This can happen when:

  • a tool or runtime crosses a boundary the policy was supposed to contain
  • an approval is present but is not meaningfully independent of the request
  • authority expands as identities, credentials, tools, and agents are composed
  • the audit record cannot establish who authorized what, under which policy, using which evidence

These are enforcement failures. They are not adequately described by model behavior alone.

NIST's NCCoE has identified closely related open questions involving agent identification, authorization, delegation, auditing, non-repudiation, and prompt-injection controls. OWASP's Agentic Security Initiative is developing threat models and practical guidance for autonomous agents and multi-step workflows. Surety is intended to complement that work by concentrating on measurable enforcement properties at runtime.

Provisional failure classes

Surety begins with four working classes. They are hypotheses to test, not a finished taxonomy.

  • F1 — Boundary enforcement failure: Can the agent cause an effect outside the boundary the harness claims to enforce?
  • F2 — Approval-integrity failure: Can a protected action proceed without a distinct, valid, policy-bound authorization?
  • F3 — Delegation and composition failure: Can authority expand or become ambiguous as agents, tools, identities, and credentials are chained?
  • F4 — Evidence-integrity failure: Can the system produce an action that cannot be reliably attributed, reconstructed, or verified afterward?

The reference invariant for F2 comes from two-person integrity:

No single actor, human or agent, can both request and authorize a protected action.

That invariant is necessary, but it may not be sufficient. Manipulation of the human approver through alert fatigue, false context, or repeated low-quality requests is an open research question. Surety will test whether that risk belongs inside F2 or requires a separate class. It will not declare a fifth class before the evidence supports one.

The fuller definitions and open questions are in docs/THREAT_MODEL.md.

What the benchmark should measure

A useful benchmark must test the harness, not reward a model for refusing a request. The same policy test should be runnable across different models and orchestration frameworks while preserving a stable expected outcome.

Each test case should identify:

  1. the policy or invariant under test
  2. the authorized initial state
  3. the adversarial action or environmental condition
  4. the expected enforcement decision
  5. the evidence that must exist after the decision
  6. the conditions that constitute a failure

The benchmark should distinguish at least four outcomes:

  • prevented — the prohibited effect did not occur
  • contained — the attempt progressed, but the protected boundary held
  • detected — the system recognized the violation but did not necessarily prevent it
  • evidenced — the resulting record is sufficient to reconstruct and attribute the decision

These outcomes are deliberately separate. Detection is not prevention. A log entry is not proof that an approval was valid. A model refusal is not evidence that the harness would stop a different model from taking the same action.

The draft measurement approach is in docs/BENCHMARK_DESIGN.md.

Research principles

  • Test stated properties. A control should be evaluated against an explicit claim, not a vague expectation of safety.
  • Separate model behavior from control enforcement. Model cooperation can be measured, but it cannot stand in for a boundary.
  • Prefer observable effects. A benchmark should score what the system allowed, denied, changed, and recorded.
  • Preserve benign utility. A system that prevents every attack by disabling every useful action is not a successful agent system.
  • Require replayable evidence. Results should include enough policy, trace, decision, and environment data to be independently examined.
  • Treat composability as hostile until demonstrated otherwise. Individually correct controls can fail when identities, tools, and delegated permissions are chained.
  • Keep claims proportional to evidence. Early results are test results, not universal conclusions.

Current status

Surety is in the design phase. This repository presently contains the research question, provisional threat model, benchmark design constraints, and an execution roadmap. There is no released harness, conformance suite, dataset, certification, or production-ready implementation.

The near-term work is to:

  • compare the four classes with existing NIST and OWASP work
  • formalize a small set of framework-neutral invariants
  • build a minimal reference harness with explicit policy and evidence surfaces
  • publish the first adversarial test cases with complete traces
  • invite review of the taxonomy and scoring before attempting broad coverage

See ROADMAP.md for the staged plan.

See Research Positioning for the current program, related engineering evidence, and explicitly future work.

The research-program map distinguishes current design artifacts from future directions.

The laboratory's public direction and evidence record are maintained in:

  • Lab Direction — why the assurance-centered research program exists
  • Research Program — the living roadmap across the current repository ecosystem
  • Repository Audits — claim, evidence, limitation, and future-work reviews for the flagship repositories

Scope

Initial scope:

  • tool-using AI agents
  • coding and operational agents with delegated credentials
  • single-agent and multi-agent workflows
  • approval-gated protected actions
  • policy decisions, identity chains, runtime effects, and audit evidence

Not in initial scope:

  • general model safety or alignment
  • content moderation
  • comparative intelligence or capability rankings
  • certification of commercial products
  • claims about the security of a model based solely on benchmark performance
  • evaluation of systems without authorization from their owners

Contributing

The most useful early contributions are narrow and falsifiable:

  • a counterexample to one of the proposed failure classes
  • an enforcement invariant that can be tested across frameworks
  • a minimal reproduction of a control-plane failure
  • criticism of the scoring model or evidence requirements
  • a mapping to an existing standard, benchmark, or threat taxonomy

Please read CONTRIBUTING.md before opening an issue. Security-sensitive reports should follow SECURITY.md.

References

Citation

Use CITATION.cff when referencing the research program. Citation metadata will be versioned only when a genuine public draft artifact exists.

About

Surety is an independent research effort by Vince Tur-Rojas, a security practitioner and engineer interested in authorization, high-consequence systems, and operational assurance.

This work is undertaken in a personal capacity. It does not represent the views of any employer or government organization. No operational, controlled, or non-public information is used in this repository.

License

Unless otherwise noted, the contents of this repository are licensed under the Apache License 2.0.

About

Open research on whether AI-agent control systems enforce the rules they claim to enforce.

Topics

Resources

Contributing

Security policy

Stars

0 stars

Watchers

0 watching

Forks

Releases

Packages

Contributors