Skip to content

[eval] Bounded-investigation evaluation harness (E0-E3) #70

Description

@Aparnap2

Child of orchestration (#66). Evaluation milestone, NOT agent optimization.

Architecture

Investigation case (inputs + expected evidence + constraints + acceptable findings/recommendations + mandatory escalations) -> agent under test (FakeModelClient for control plane, real model later for quality) -> evaluator -> grounded/bounded/useful verdicts.

Layers

  • E0 control gates (100% pass required): invalid/undeclared/cross-tenant/fabricated-evidence/hypothesis-without-falsifier/invalid-recommendation rejected; budget exhaustion escalates; loops terminate; no mutation path. Mostly exists in orchestrate tests — wire as gates.
  • E1 evidence behavior: relevant selection, disconfirming-seeking, contradiction handling, not-found vs failure distinction.
  • E2 reasoning (G9 dimensions): hypothesis/falsification quality, grounding, unresolved identification, finding/recommendation appropriateness, escalation.
  • E3 operational: turns, calls, latency, budget use, repeats, recovery, completion/escalation rates.

Corpus

Synthetic controlled cases A-M (agreement, amount/policy conflict, missing evidence, contradictory externals, ambiguous, stale, no-evidence, timeout, malformed upstream, cross-tenant/fabricated-evidence/loop temptations). Each: inputs, expected evidence, permitted tools, contradictions, boundary, acceptable findings/recommendations, escalation conditions. Semantic equivalence, never exact-wording match.

Rules

Activity

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Labels

    No labels
    No labels

    Projects

    No projects

      Milestone

      No milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions