Skip to content
Closed
Show file tree
Hide file tree
Changes from all commits
Commits
Show all changes
34 commits
Select commit Hold shift + click to select a range
bffdc04
Add clean no-skill host eval mode
GeoGeekLab Sep 23, 2026
8c80e85
Test clean no-skill adapter behavior
GeoGeekLab Sep 23, 2026
3c2f407
Support repeated eval runs and timing evidence
GeoGeekLab Sep 23, 2026
b69e203
Test repeated eval metadata
GeoGeekLab Sep 23, 2026
4341d2b
Add Skill A/B eval comparison report
GeoGeekLab Sep 23, 2026
8144e02
Test Skill A/B result comparison
GeoGeekLab Sep 23, 2026
b03462f
Document repeated eval run index
GeoGeekLab Sep 23, 2026
434a0f6
Document Skill A/B evaluation workflow
GeoGeekLab Sep 23, 2026
1e94c9a
Require paired repetitions in A/B comparisons
GeoGeekLab Sep 23, 2026
9214276
Test paired A/B repetition requirements
GeoGeekLab Sep 23, 2026
cddaf88
Add paired counterbalanced A/B eval runner
GeoGeekLab Sep 23, 2026
ec6c4b2
Keep paired runner environment handling self-contained
GeoGeekLab Sep 23, 2026
935b915
Test paired counterbalanced A/B runner
GeoGeekLab Sep 23, 2026
e10d1da
Add manual real-host Skill A/B workflow
GeoGeekLab Sep 23, 2026
c26fe7d
Harden real-host workflow inputs and diagnostics
GeoGeekLab Sep 23, 2026
f9ba730
Document paired real-host A/B phase
GeoGeekLab Sep 23, 2026
69b2bbd
Require A/B evaluation tooling in repository validation
GeoGeekLab Sep 23, 2026
6c6a2ab
Revert real-host experiment additions in scripts/compare_eval_results.py
GeoGeekLab Sep 23, 2026
e1a2cae
Revert real-host experiment additions in tests/test_compare_eval_resu…
GeoGeekLab Sep 23, 2026
f3e1d65
Revert real-host experiment additions in evals/README.md
GeoGeekLab Sep 23, 2026
424c520
Revert real-host experiment additions in scripts/validate_skill.py
GeoGeekLab Sep 23, 2026
1f4e568
Remove real-host experiment file .github/workflows/real-host-ab.yml
GeoGeekLab Sep 23, 2026
4cc5330
Remove real-host experiment file scripts/run_ab_evals.py
GeoGeekLab Sep 23, 2026
7041e3a
Remove real-host experiment file tests/test_run_ab_evals.py
GeoGeekLab Sep 23, 2026
ed1dbbc
Add semantic claim and generator-drift checks
GeoGeekLab Sep 23, 2026
cb4f43b
Document stronger hidden eval checks
GeoGeekLab Sep 23, 2026
8bc2c82
Strengthen evals for second-order engineering risks
GeoGeekLab Sep 23, 2026
9d80768
Test second-order eval check semantics
GeoGeekLab Sep 23, 2026
114b574
Document strengthened second-order eval coverage
GeoGeekLab Sep 23, 2026
7341a26
Fix negation-aware final claim matching
GeoGeekLab Sep 23, 2026
a155ff1
Validate hidden inline Python checks without executing them
GeoGeekLab Sep 23, 2026
ce54e91
Test hidden-check validation and negation edge case
GeoGeekLab Sep 23, 2026
2dc81d1
Regression-test flaky fixture module isolation
GeoGeekLab Sep 23, 2026
ab882e3
Sharpen README for GitHub engineering style
GeoGeekLab Sep 23, 2026
File filter

Filter by extension

Filter by extension

Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
479 changes: 330 additions & 149 deletions README.md

Large diffs are not rendered by default.

61 changes: 52 additions & 9 deletions evals/README.md
Original file line number Diff line number Diff line change
Expand Up @@ -21,7 +21,7 @@ A fake or deterministic adapter used by unit tests proves only that the harness

## Case structure

`cases.json` contains 14 executable scenarios. Each case defines:
`cases.json` contains 17 executable scenarios. Each case defines:

- `id`: stable kebab-case identifier,
- `task`: the instruction exposed to the agent,
Expand All @@ -44,8 +44,9 @@ The runner currently supports:
- file existence and absence,
- unchanged-file assertions,
- required or forbidden file content,
- final-output term assertions,
- executable commands with optional repetition.
- final-output term assertions, including negation-aware forbidden-claim checks,
- executable commands with optional repetition,
- idempotence checks that run a generator or formatter and fail if it produces a diff.

Executable commands are useful for regression tests, compatibility tests, generators, and repeated flaky-test checks. They are not treated as inherently safe.

Expand Down Expand Up @@ -128,6 +129,47 @@ Adapter unit tests validate command construction and staged-skill handling. Thos

Each case receives a fresh runtime Skill staging directory. The harness hashes it before and after the agent exits. Any mutation produces a failing `skill_payload_integrity` check, so one case cannot rewrite the Skill used by later cases.

### Compare Skill vs no-Skill behavior

The native adapter can run the same task without exposing the staged Skill. The baseline is created by removing the Skill from the host environment, not by adding a prompt that tells the model to ignore it.

Keep the host, explicit model ID, case set, credentials, execution flags, and task text identical between conditions. Use `--repeat` when you need repeated independent trials; every repetition receives a fresh fixture workspace.

Codex example:

```bash
python scripts/run_evals.py \
--agent-command '{python} {repo}/scripts/host_eval_adapter.py codex --model MODEL --skill-mode disabled' \
--adapter-label codex-MODEL-no-skill \
--pass-env CODEX_API_KEY \
--repeat 5 \
--allow-workspace-execution \
--output eval-results/codex-MODEL-no-skill.json

python scripts/run_evals.py \
--agent-command '{python} {repo}/scripts/host_eval_adapter.py codex --model MODEL --skill-mode enabled' \
--adapter-label codex-MODEL-skill \
--pass-env CODEX_API_KEY \
--repeat 5 \
--allow-workspace-execution \
--output eval-results/codex-MODEL-skill.json
```

Claude Code uses the same `--skill-mode disabled|enabled` switch on `host_eval_adapter.py`.

Compare the two reports:

```bash
python scripts/compare_eval_results.py \
eval-results/codex-MODEL-no-skill.json \
eval-results/codex-MODEL-skill.json \
--output eval-results/codex-MODEL-comparison.md
```

The comparison reports deterministic case pass rates, deterministic check-type pass rates, mean changed-file counts, and mean agent wall-clock time. It deliberately does not convert the qualitative `must_do` / `must_not_do` rubric into an automatic score. The infrastructure-only `skill_payload_integrity` check is excluded from comparative check rates.

For publishable evidence, run both conditions close enough together to reduce host/model drift, retain the raw JSON reports, and review qualitative rubric items separately. Token usage is not currently normalized across host CLIs, so wall-clock time is the portable cost signal recorded by the harness.

See [host compatibility](../docs/compatibility.md) for the dated vendor documentation basis.

## Execution boundary
Expand Down Expand Up @@ -160,15 +202,16 @@ The current scenarios cover:
- public-contract migration,
- incorrect abstraction pressure,
- honest verification reporting,
- path traversal boundaries,
- path traversal, symlink, TOCTOU, and tenant-authorization boundaries,
- speculative performance optimization,
- bounded concurrency and ordering,
- bounded concurrency, ordering, and failure cleanup,
- behavior-preserving refactoring,
- unnecessary dependency pressure,
- flaky-test repair,
- mixed-version database migration,
- generated-code source-of-truth changes,
- swallowed errors,
- uncontrolled scope expansion.
- mixed-version and rollback-safe database migration,
- generated-code source-of-truth changes and generator-drift detection,
- swallowed errors and public error-contract compatibility,
- blocked verification evidence,
- uncontrolled scope expansion in the presence of a known unrelated failure.

When the skill contract changes, add or strengthen a case that can distinguish the old behavior from the intended behavior. Prefer executable invariants over prose-only expectations when the behavior can be measured.
Loading
Loading