Skip to content

Add Skill vs no-Skill behavioral eval comparison - #21

Closed
GeoGeekLab wants to merge 34 commits into
mainfrom
eval/skill-ab-comparison
Closed

GeoGeekLab wants to merge 34 commits into
mainfrom
eval/skill-ab-comparison

Conversation

@GeoGeekLab

@GeoGeekLab GeoGeekLab commented Sep 23, 2026 •

Copy link
Copy Markdown
Owner

Summary

This PR adds controlled Skill-vs-no-Skill evaluation support and strengthens the behavioral suite around second-order engineering risks.

A/B harness

  • add --skill-mode disabled|enabled to native Codex and Claude Code adapters
  • add repeated runs with fresh workspaces, run indexes, and agent wall-clock timing
  • add a case-by-case comparison report without collapsing engineering quality into one score
  • keep qualitative must_do / must_not_do criteria outside automatic scoring

Eval correctness fixes

  • fix flaky-test by renaming fixture module token.py to token_value.py, avoiding collision with Python's stdlib token module
  • add a regression test that materializes the fixture and imports the production module from the isolated workspace
  • replace the verification false-positive substring check with final_not_claim_any
    • negated statements such as “not fully verified” are allowed
    • positive claims such as “fully verified” still fail
    • “not only fully verified …” is treated as a positive claim
  • add command_no_changes so generator checks can require a true fixed point instead of silently repairing stale generated output
  • compile hidden python -c checks during --validate-only without executing fixture code

Stronger second-order cases

The suite grows from 14 to 17 scenarios.

Existing scenarios now include stronger hidden evidence for:

  • security-boundary: symlink escape
  • concurrent-worker: failure propagation and pending-work cancellation
  • database-migration-rollout: backfill, old-writer/new-reader compatibility, and rollback-safe dual writes
  • generated-code-change: generator drift / idempotence
  • verification-evidence: execution markers proving local and blocked checks were actually attempted
  • scope-expansion: a known unrelated failing regression that must remain out of scope

New scenarios:

  • security-toctou: file replacement between validation and open
  • authorization-boundary: cross-tenant object access
  • api-error-compatibility: exception type, code, and message compatibility

Most of these second-order assertions live in hidden deterministic checks rather than visible fixture tests, so the fixture does not simply reveal the intended answer.

Validation

Latest PR head passes:

  • quality: Python 3.10
  • quality: Python 3.12
  • quality: Python 3.14
  • package: success

Repository validation, compilation, unit tests, eval-schema validation, package checks, and release invariants all pass.

The negation-aware claim check is intentionally a small deterministic token-window rule, not an LLM semantic judge.

@GeoGeekLab GeoGeekLab closed this Sep 23, 2026
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant