Skip to content

feat: discriminative code-gen scoring — execution scorer + 108-case battery - #1

Draft
Ryfter wants to merge 39 commits into
masterfrom
feat/codegen-execution-scorer
Draft

Ryfter wants to merge 39 commits into
masterfrom
feat/codegen-execution-scorer

Conversation

@Ryfter

@Ryfter Ryfter commented Jul 4, 2026 •

Copy link
Copy Markdown
Owner

Summary

Item 1 of the discriminative scoring v2 roadmap: make code-gen actually measure capability. Started as an execution scorer; now also carries the battery that makes the scorer worth running.

compilable-code only checked that output parses, so garbage that compiles scored 1.0 and nearly every model read 1.0. The cases were memorised classics on top of that.

What changed

Scorer

  • code-exec: runs the candidate in an isolated subprocess against hidden asserts; score = fraction passed, not binary.
  • Cross-platform sandbox teardown. The original timeout path used os.killpg/SIGKILL, both POSIX-only — on Windows the runaway child survived and TemporaryDirectory cleanup died with WinError 32. Timeouts are the characteristic weak-model failure, so this broke the most-exercised path on the box Gauntlet actually runs on. Now taskkill /F /T vs killpg, always reaped, pipes closed.
  • Structured failure_mode: none / no_code_emitted / syntax_error / runtime_exception / wrong_answer / timeout / harness_error. Aggregated per cell. This is the signal that distinguishes a capability gap from a defect in plausible code — only the second is worth scaffolding (D-2026-06-30c).
  • AST constraint checking (no-imports, single-expression, must-use-generator, …) reported separately from correctness, isolating instruction-following from algorithmic skill.
  • Per-tier scorecard profile. The shape beats the mean: clearing T1–T2 then collapsing at T3 is a different proposition from scoring evenly.

Battery — 108 cases, 12 dimensions

tier T1 16 (15%) · T2 28 (26%) · T3 38 (35%) · T4 26 (24%)
dimensions adversarial-correctness, stdlib-api-use, class-level-stateful, multi-function, execution-prediction, complexity-constrained, bug-fix, surface-constraints, robustness-contracts, behavior-preserving-refactor, test-authoring, data-text-munging

The ladder is deliberate. The fleet is 1B–30B local models: all-1.0 is non-discriminative, but so is all-0.0. T3 does most of the discriminating; T4 is headroom so the battery doesn't saturate as models improve.

Dimensions are adapted from EvalPlus/HumanEval+, BigCodeBench, ClassEval, CRUXEval, EffiBench/BigO(Bench), DS-1000, SWE-bench and TREAT.

Anti-contamination. Retired fizzbuzz, palindrome, binary-search, lru-cache, csv-parse, plus four classic-shaped cases from the first pass (run-length-encode, matrix-transpose, safe-divide, factorial-strict). Files remain in the tree; they are out of the battery. New cases are original problems or familiar shapes with a twist a memorised answer gets wrong.

Every case proves itself. Each ships a reference solution (must score exactly 1.0) and a subtly-wrong solution (must score < 1.0), enforced for all 108 by tests/test_case_validation.py. The gate rejected 5 authored cases: three whose reference could not satisfy its own asserts, one whose wrong solution still scored 1.0, one duplicate-weak. An unsatisfiable case scores every model 0.0 and a toothless one scores every model 1.0 — both look like data and are noise.

Discrimination evidence

Five synthetic outputs against all 108 cases:

excellent (reference)      mean=1.000   none=62
subtly wrong               mean=0.635   wrong_answer=62
prose, no code             mean=0.000   no_code_emitted=62
syntax error               mean=0.000   syntax_error=62
infinite loop              mean=0.000   timeout=62

Under compilable-code, the first two both scored 1.0. Subtly-wrong landing at 0.635 is the point: partial credit, not a cliff.

That sweep also caught a real bug — prose replies were being filed as syntax_error (true of the bytes, wrong about what happened), since the check only caught empty output. Fixed, with a test pinning the boundary so a genuine-but-malformed attempt still reads as syntax_error.

Tests

588 passed, 8 skipped, 4 deselected in 73.41s

Skips are test-authoring cases exempted from the prompt-leak check: there the prompt legitimately contains the implementation under test, so overlap is the task, not a leak.

Cost

108 cases × 9 models ≈ 970 cells per full run. Sandbox overhead ~0.3–0.5s per case on top of inference.

Not done here

  • Anti-cheat/integrity layer — subprocess isolation is not a security boundary. Absolute-path filesystem reads and direct network syscalls are still open, so a candidate could in principle read its own hidden tests. Tracked separately.
  • Roadmap items 2–5 (v2 spec, raw-vs-scaffolded, pipeline economics, ToC scheduler) need local inference and design sign-off.

🤖 Generated with Claude Code


Integrity layer (anti-cheat)

A benchmark that can be shortcut measures nothing. The sandbox was subprocess-only — absolute-path reads and network syscalls were unrestricted — so a candidate could read its own hidden tests and score 1.0 without demonstrating anything.

vector control
A prompt leakage line-level guard (≥40 chars; shorter runs collide constantly on ordinary Python). test-authoring exempt by design
B filesystem in-sandbox patching of open/io.open/os.open/listdir/scandir against the benchmark tree and the sandbox's own hidden_tests.py copy; sys.path stripped; process creation blocked
C network socket entry points refuse and record
D tool use client.py refuses tools/functions/web_search/… and inspects reply shape — it is the only HTTP component, so nothing bypasses it
E canaries deterministic per-case marker in every hidden-test file; an echo in model output marks the cell unscored, never 0

Two holes found by attack, not by reasoning

Both succeeded against the first version of the guard:

  • os.system("cp hidden_tests.py leaked.txt"), then read the copy. The copy had an allowed name, so the open() patch never saw it. Fixed by blocking process creation as a class rather than chasing filenames.
  • gc.get_objects() → the hidden module's __dict__. The grader necessarily shares the candidate's process, so the expected values are reachable. gc.get_objects, gc.get_referrers and sys._getframe are now blocked.

Results

red team (16 attacks, authored blind by an outside model):
  0/16 obtained the answers,  9 recorded as explicit violations

white box (4 attacks, written with knowledge of runner internals):
  0/4  obtained the answers,  4 recorded as explicit violations

Both numbers are reported because "obtained nothing" and "was blocked" are different claims — a cheat that never worked is not evidence the guard stopped it.

The corpus is committed under tests/redteam/ so future changes can be re-tested against it.

What this is not

docs/2026-07-25-benchmark-integrity-threat-model.md states the limits plainly rather than implying coverage that does not exist:

  • A Python-level guard is defeatable by ctypes, re-exec, or raw syscalls. It blocks accidental and casual cheating — the realistic failure for cooperative local models — and is not a security boundary.
  • In-process grading is structurally reachable. Closing that properly means grading out-of-process, which is incompatible with calling arbitrary candidate functions. Flagged for revisit if Gauntlet is ever used to rank models with something at stake.
  • A gateway performing retrieval silently, declaring nothing in request or response, is undetectable from here.

Fail closed: a control that cannot be enabled marks affected cells unscored with the reason recorded — never silently scored as though it were active. Violations likewise yield unscored, not 0.0: zero is a claim about the model's ability, unscored is an admission that we do not know.

Tests

611 passed, 8 skipped, 4 deselected in 81.51s

claude and others added 10 commits July 4, 2026 11:10
Tracked, ordered roadmap for the v2 initiative seeded in
docs/2026-06-30-discriminative-scoring-v2-seed.md: code-gen execution scorer
first (contained, no local inference, unblocks every downstream code number),
then the full spec (gates raw-vs-scaffolded, pipeline economics, and the
ToC-aware scheduler, which all need the multi-axis scoring contract nailed
down first). Carries forward the selective-offload guardrail so item 3
exposes the working-memory-vs-capability-gap distinction instead of papering
over it.
compilable_code_match only checks that output parses (compile(..., "exec")),
so garbage that compiles scores 1.0 and code-gen is non-discriminative -
nearly every model hits 1.0, including on toy cases memorized by 1B models.

Add a code-exec scoring method alongside compilable-code (kept intact for
backward compat, selectable per case): it runs generated code in an isolated
subprocess (fresh process, own process group so a timeout also kills any
children, scratch tempdir as cwd - not the repo tree, minimal env with no
inherited proxy/API-key vars) against a hidden, maintainer-authored assert
suite (tests_file: a plain Python file defining check(ns) -> list[bool]).
Score is the fraction of hidden asserts passed; passed requires all of them.
Syntax errors, exceptions, and timeouts all score low and never crash the
runner. A broken hidden-test harness (our bug, not the candidate's) is
recorded unscored, never silently 0, per the scoring-honesty invariant.

Adds 6 harder code-gen cases (run-length-encode, matrix-transpose,
safe-divide, factorial-strict, inventory-tracker, prime-pair) covering edge
cases, a stateful class, and a multi-function ask - not memorized classics.
Each hidden test suite was sanity-checked against a reference-correct and a
deliberately-wrong solution.

Item 1 of the discriminative-scoring-v2 roadmap
(docs/superpowers/plans/2026-07-04-discriminative-scoring-v2-plan.md).

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01TvMbRwJkGqEreYKCNc8hTq
The timeout path used os.killpg/SIGKILL, both POSIX-only, so on Windows the
runaway child was never killed and TemporaryDirectory cleanup died on an open
handle (WinError 32). Timeouts are the common small-model failure, so this hit
the exact path a real battery run exercises most.

- _kill_tree(): taskkill /F /T on Windows, killpg on POSIX; always reaps and
  closes pipes before the scratch dir is removed.
- _spawn_kwargs(): CREATE_NEW_PROCESS_GROUP vs start_new_session.
- _sandbox_env(): Windows needs SystemRoot or the interpreter will not start.
- ExecutionResult.failure_mode: structured reason (syntax_error /
  runtime_exception / wrong_answer / timeout / no_code_emitted /
  harness_error) so the scorecard can aggregate why a model failed instead of
  parsing prose. Feeds the selective-offload principle (D-2026-06-30c).
- no_code_emitted is now distinct from broken code.
- Case.tier (T1-T4) + Case.dimension for the difficulty ladder.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
- constraints.py: declarative AST checks (no-imports, single-expression,
  must-use-generator, no-sorted, no-loops, no-recursion, ...). Reported
  separately from correctness so an instruction-following miss is not
  conflated with a capability gap -- the distinction the selective-offload
  principle turns on (D-2026-06-30c). Unparseable source reports no violation:
  that is a syntax_error, and double-counting it would misattribute one
  failure as two.
- test_case_validation.py: every registered case must prove itself. Reference
  solution must score exactly 1.0 (asserts satisfiable) and a deliberately
  wrong solution must score below 1.0 (asserts have teeth). Also asserts
  hidden test content never leaks into the prompt. A broken case now fails CI
  instead of quietly skewing a scorecard.
- First 4 execution-scored cases (adversarial-correctness, T1-T4) plus a
  registry the battery YAML and meta-test are generated from.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
…ation

Documents the design the expanded code-gen battery is built on: the T1-T4 tier
ladder (all-1.0 and all-0.0 are equally non-discriminative for a 1B-30B fleet),
the eleven capability dimensions adapted from EvalPlus/BigCodeBench/ClassEval/
CRUXEval/EffiBench/DS-1000/SWE-bench/TREAT, the failure-mode table, the
anti-contamination rule, and the prove-itself contract every case must satisfy.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
…ons)

Grok-authored against a strict contract, each machine-validated before landing:
reference solution must score 1.0 and a subtly-wrong solution must score <1.0.
One case (excel-col-range-union) was rejected by that gate because its own
reference solution scored 0.9 -- an unsatisfiable case that would have capped
every model.

Dimensions so far: adversarial-correctness, stdlib-api-use,
class-level-stateful, multi-function. Tier spread T1 21% / T2 26% / T3 32% /
T4 21%.

The contaminated classics (fizzbuzz, palindrome, binary-search, lru-cache,
csv-parse) are dropped from the battery; the four classic-shaped code-exec
cases from the first pass (run-length-encode, matrix-transpose, safe-divide,
factorial-strict) are likewise not carried over. Files remain in the tree.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Adds execution-prediction (CRUXEval-style: the model traces a function and
predicts outputs rather than writing code -- reasoning, not generation, scored
fractionally via ten ANSWER_n assignments) and test-authoring (the model writes
tests against an implementation with a planted bug; scored on whether its tests
actually catch it).

test-authoring is exempted from the prompt-leak check in both the ingest
validator and the meta-test: there the prompt legitimately contains the
implementation under test, so overlap with the hidden tests is the task rather
than a leak. The bug being hunted is still never revealed.

Tier spread T1 15% / T2 26% / T3 35% / T4 24% -- on target for a ladder that
discriminates across a 1B-30B fleet instead of saturating at either end.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Threads tier/dimension/failure_mode from the case through CaseResult into the
Cell, and aggregates two new views:

- quality_by_tier: mean per T1-T4. The shape carries more than the mean -
  clearing T1-T2 then collapsing at T3 is a different proposition from scoring
  evenly across all four. Unscored cases are excluded, not counted as 0.
- failure_modes: how many cases failed each way. This is the signal Baton needs
  to tell a capability gap (no_code_emitted / syntax_error - the model cannot
  engage) from a defect in plausible code (wrong_answer). Only the latter is
  worth scaffolding (D-2026-06-30c).

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
…error

A discrimination sweep over all 62 cases showed every prose reply filed as
syntax_error. True of the bytes, wrong about what happened: the previous check
only caught *empty* output, so 'I'd be happy to help! Could you clarify...'
reached the compiler and failed there.

_looks_like_code() now asks whether the model attempted Python at all --
anything that parses is code, anything that does not is only a syntax error if
it at least reaches for the language. Deliberately generous so a genuine but
malformed attempt still reads as syntax_error rather than silence; a test
pins that boundary.

This matters because the two failures want different responses: a model that
never engages is a capability gap, one that writes broken code is not.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Wave 2 tops up every thin dimension. Final shape:

  tier   T1 16 (15%)  T2 28 (26%)  T3 38 (35%)  T4 26 (24%)
  dims   class-level-stateful 12, execution-prediction 12, multi-function 11,
         surface-constraints 10, adversarial-correctness/behavior-preserving-
         refactor/bug-fix/data-text-munging/robustness-contracts/stdlib-api-use
         8 each, test-authoring 8, complexity-constrained 7

Three more cases were rejected by the validation gate: two whose reference
solution could not satisfy its own asserts, and one whose deliberately-wrong
solution still scored 1.0 (asserts too weak to discriminate). Both directions
of the gate are doing work.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
@Ryfter Ryfter changed the title feat: execution-based code-gen scorer (discriminative scoring v2, item 1) feat: discriminative code-gen scoring — execution scorer + 108-case battery Jul 25, 2026
Ryfter and others added 19 commits July 25, 2026 15:40
A benchmark that can be shortcut measures nothing. The sandbox was
subprocess-only: absolute-path reads and network syscalls were unrestricted, so
a candidate could read its own hidden tests and score 1.0 without demonstrating
anything.

Controls, by vector:
- A prompt leakage: line-level n-gram guard at >=40 chars (shorter runs collide
  constantly on ordinary Python). test-authoring exempt by design.
- B filesystem: in-sandbox guard patching open/io.open/os.open/listdir/scandir
  against the benchmark tree AND the sandbox's own hidden_tests.py copy, plus
  sys.path stripping. Process creation blocked as a class.
- C network: socket entry points refuse and record.
- D tool use: client.py refuses requests carrying tools/functions/web_search
  and inspects reply shape. It is the only HTTP component, so nothing bypasses it.
- E canaries: deterministic per-case marker in every hidden-test file; an echo
  in model output marks the cell unscored, never 0.

Two holes were found by attack, not by reasoning, and both are now detected:
  os.system("cp hidden_tests.py leaked.txt") then reading the copy -- the copy
  had an allowed name, so the open() patch never saw it. Blocking process
  creation closes the class rather than chasing filenames.
  gc.get_objects() reaching the hidden module's __dict__ for the expected
  values, since the grader necessarily shares the candidate's process.

Results: red team 0/16 obtained answers (9 recorded as violations); white-box
0/4 (4 recorded). Both numbers reported because "obtained nothing" and "was
blocked" are different claims.

The threat model states plainly that this is not a security boundary: a
Python-level guard is defeatable via ctypes or re-exec, and in-process grading
is structurally reachable. It blocks accidental and casual cheating, which is
the actual threat model for cooperative local models.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
These scripts built the 108-case battery and lived only in a session
scratchpad, so the ability to EXTEND the battery would have been lost while the
battery itself survived.

scripts/battery-authoring/ holds the draft -> validate -> regenerate ->
verify-discrimination flow, plus the authoring contract given to the drafting
model. The validation gate is what makes delegating case drafting to a cheap
model safe: a case cannot land unless its reference solution scores 1.0 and its
wrong solution scores below it, so quality never rests on trusting the drafter.

Also moves the red-team cheat prompt next to its corpus.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Supersedes the stale parts of the 2026-07-04 code-exec record (battery "grows to
11 cases"; sandbox "does not block network syscalls or filesystem access") --
both are now out of date.

Records the six decisions behind the 108-case battery and integrity layer, each
with the alternative rejected: the difficulty ladder over merely-harder cases,
twelve dimensions for breadth, retiring the contaminated classics, the
prove-itself gate (which is what made delegating case drafting safe), structured
failure modes, and the honestly-scoped integrity layer.

Propagates in-flight state to AGENTS.md and GEMINI.md so other agents know PR #1
is open and unmerged, that the battery YAML is generated rather than hand-edited,
and that changes must be tested on Windows -- an earlier cloud-authored change
shipped POSIX-only teardown that broke every timeout on Kevin's box.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
A cell is a summary, and a summary only answers the questions it was designed
for. Two gaps made the 108-case battery unable to answer the question it exists
to answer:

- No per-dimension breakdown. "Best at code-gen" is not actionable; "best at
  bug-fix, weak on multi-function" is, and that is the axis Baton routes on.
  Adds `quality_by_dimension` alongside the existing tier profile, with the
  same rule: unscored cases are excluded, never counted as 0.
- Per-case results were discarded once aggregated. A full fleet run costs hours
  of GPU time, so any new question -- which cases did every model miss? is this
  dimension too hard? -- cost another whole run. Adds `cases.jsonl` beside
  `cells.jsonl`, one row per case, checkpointed before the cell so a crash
  between the two loses only what resume rebuilds.

`score` stays null on the way to disk for unscored cases; a zero is a claim the
model failed, a null is an admission we do not know.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
`scripts/analyze_fleet.py` turns a run directory into the report the benchmark
exists to produce: per-capability leaderboard, best model per dimension,
difficulty-curve shape, failure-mode read, and a throughput-adjusted ranking.

Drafted by a fleet model and gated on `tests/test_analyze_fleet.py`, per the
standing rule that delegated output does not land on trust. The gate earned its
keep -- three defects it caught:

- The report crashed on write. It uses em dashes and a Unicode minus, and a
  Windows console is cp1252. Passed on Linux, fatal on the box Gauntlet runs on.
- The throughput ranking collapsed a model's several capability cells by keeping
  its best one, reporting every model at its strongest job as if it were the
  model's number. Replaced with a case-weighted mean so a 108-case battery
  outweighs an 8-case one.
- The shape classifier called a sawtooth (0.2 / 0.9 / 0.3 / 0.85) a `cliff`
  because it only looked for one big adjacent drop. A cliff means "holds up,
  then falls off and stays down"; a profile that climbs back is noise. Calling
  it a cliff would tell Baton the model is reliable up to a tier when it isn't.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
`_strip_fences` only unwrapped a reply that *started* with a fence. Models do
not answer that way -- they say "Here's the solution:", fence the code, and sign
off. Every one of those replies went to `ast.parse` with the prose and the
backticks still attached and came back `syntax_error`.

Measured on llama-3.2-1b over the 108-case battery:

    syntax_error   76 -> 14      (62 cases were mis-scored)
    wrong_answer   27 -> 78
    quality      0.139 -> 0.338

The number was wrong, but the reading was worse. A syntax_error-dominant
profile says "capability gap -- the model cannot engage with the task"; a
wrong_answer-dominant one says "scaffolding candidate -- plausible code with a
defect". Only the second is worth scaffolding (D-2026-06-30c). The bug inverted
exactly the routing call the battery exists to inform, and it did it uniformly
against the chattiest models rather than the worst ones -- a bias that reads as
a finding.

`extract_code` takes the longest fenced block anywhere in the reply, tolerates
tilde fences and a block truncated by the token limit, and returns unfenced
text untouched so `_looks_like_code` can still attribute prose as
`no_code_emitted` rather than this inventing a block the model never wrote.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Adds a third easy-to-violate rule: a scorer bug and a weak model produce identical-looking scorecards. The fence-extraction bug read as a capability gap in every symptom it produced. Records the deliberate asymmetry too -- only code-exec extracts fenced code, because the other batteries ask for bare output and prose there is a real instruction-following failure.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
The first full fleet run scored seven of nine models at ~0.00 on code-gen, almost entirely no_code_emitted. They had not failed to write code: they are reasoning models that spent the whole 512-token budget thinking and were cut off before the answer. gemma-4-12b, scored 0.01, writes correct code at 4096 tokens (finish_reason stop, 859 chars of content). The benchmark was measuring its own configuration and reporting it as a property of the model.

Three changes. The client now surfaces finish_reason. A truncated reply whose failure is consistent with being cut off (no_code_emitted / syntax_error) becomes unscored with failure_mode=truncated, per the scoring-honesty invariant -- we cannot tell whether the model could not do it or was not allowed to finish, and unscored is the only honest answer. Batteries carry their own max_tokens, because 'classify this in one word' and 'write this function' do not need the same room; code-gen gets 4096.

Truncation does not soften a real failure: prose emitted within budget is a genuine miss, and a wrong_answer ran to completion so its defect is real regardless of what was cut off after. gen_yaml.py emits max_tokens too, so regenerating the battery cannot silently drop it.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
The timeout is per-chunk on an SSE stream, not per-request, so it bounds the longest silence rather than total generation. A reasoning model can think for minutes before its first content token, and 120s was cutting those off as transport errors -- which reads as an unreachable box rather than a slow model.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Gauntlet runs on Kevin's desktop, not a headless box. A backgrounded run writing to a log file is invisible: the first symptom is a loud machine and no way to tell what is causing it. He had to ask whether it was me.

Adds a status file at scorecards/.running.json (run id, pid, current model, cells done/total), written at start, updated per cell, cleared in a finally. 'gauntlet status' reads it and also lists what is resident in VRAM. It reports a stale marker from a killed run as idle rather than running, with the resume command -- a stale file that reads as RUNNING is the wrong answer to the only question the file exists to answer. The file carries no base_url or host, per the privacy invariant: knowing what is running never requires knowing where.

The run also hands the GPU back on exit, because a loaded-but-idle model costs power for nothing. Which models to free is decided by snapshot-diff: what we ran, minus what was already resident when we started. loaded_models() returns None rather than [] when it cannot tell, and an unknown snapshot unloads nothing -- leaving VRAM occupied wastes power, but evicting a model from Kevin's own session breaks his work, and the second is worse. Same shape as unscored-is-not-zero: absence of knowledge is not knowledge of absence.

Unloading shells out to the lms CLI, so the OpenAIClient-is-the-only-HTTP-component invariant is untouched. The sequencer already groups profile-outer, so a model loads once per run and all seven batteries run while it is resident; no change needed there.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Two defects found by running the thing rather than reasoning about it.

1. is_alive used os.kill(pid, 0). That is the POSIX idiom for probing a process, but on Windows CPython routes any signal other than CTRL_C_EVENT/CTRL_BREAK_EVENT straight to TerminateProcess -- so 'gauntlet status' was a way to kill the run you were asking about. It also reported a failed probe as alive, which left 'release' permanently convinced a dead run was still going. Now opens the process for query only and reads STILL_ACTIVE. Tests pin both halves: a dead pid reads dead, and probing a live one twice leaves it running.

2. Unload-on-exit lived in a finally, which does not run on SIGKILL -- and that is how an interrupted run actually ends; today's was killed three times, each time leaving ~10GB of VRAM held. The status file now persists both halves of the snapshot-diff, so 'gauntlet release' can finish the job from disk alone and still never touch a model from another session. A marker with no recorded snapshot refuses to unload rather than guessing.

Same root cause orphaned two sandbox runner.py processes from yesterday's killed discrimination sweep; reaped.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
green = inference running, yellow = a model is resident but nothing is running, red = nothing loaded. Yellow is the reason it exists: a run that failed to release its model looks exactly like an idle machine from the outside while costing power the whole time. Green tells you why the fans are up; yellow tells you something needs cleaning up.

Borderless always-on-top tkinter window (stdlib, no new dependency). Drag from anywhere, close with the corner mark or Escape, position remembered between sessions. Runs as its own process polling the run marker, so it cannot slow, block, or crash a benchmark; 'gauntlet run' spawns one by default and closes it at the end (--no-overlay to opt out), and a headless box with no display simply gets no lamp.

The state logic is a pure function with its own tests, separate from the window. Two behaviours worth keeping: a stale marker from a killed run never shows green, and an unqueryable VRAM state degrades to yellow rather than claiming the card is free -- the reassuring answer is the one most likely to be a guess.

Also adds a per-case heartbeat. A cell is one battery against one model, which for 108 code-gen cases against a verbose reasoning model runs for hours; at cell granularity alone the lamp sat unchanged for all of it, which is indistinguishable from a hung run. And a single-instance lock, because borderless windows have no title bar, so a second lamp stacks invisibly on the first and you just get a stale reading from whichever is on top.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
A model sharing the card spills layers to CPU, and the number that comes out then describes the contention rather than the model. Measured live: gemma-4-31b needs 19.89GB and glm-4.7-flash 24.61GB on a 31.8GB card, so a 14.19GB model left resident from another session forced both into offload -- 15.3 tok/s with a 30s TTFT on a 5090, and a 0.01 score that was starvation rather than incapacity.

The run now clears the local GPU first and unloads each model as soon as its last battery finishes, instead of holding it for the rest of the run. Default on (--shared-vram opts out). Keeps every model in a run comparable to every other one, and keeps power draw to what the work actually needs.

Only models on this machine are ever touched: lms ps also lists models held by linked instances on other boxes, where unloading frees nothing here and interrupts a different machine. Locality is detected by a bare 'Local' token rather than a column index, because the SIZE field contains a space and shifts everything after it.

Also stops release_after_run crediting itself with freeing models that were already gone -- lms unload exits 0 for a model that was never loaded, and a tool whose whole job is reporting machine state does not get to overstate what it did.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Three arms over the same 108 cases and the same hidden tests: raw, prompt-scaffolded, and a 3-turn decomposition. Only the prompting differs; scoring is identical, or the delta would be meaningless.

The point is the falsifiable test of D-2026-06-30c. Structured failure_mode makes it possible for the first time to predict per model x dimension which models scaffolding should help (wrong_answer-dominant) and which it should not (no_code_emitted / syntax_error). Uniform gains across both groups would disconfirm the principle, and the spec says plainly that this is a valuable outcome rather than a failed experiment.

Scaffolds are generic templates applied at run time, not per-case authored ones: Baton can only ever apply a generic scaffold automatically, so a hand-tuned decomposition would measure the author and measure something Baton can never deploy. The multi-turn arm never sees hidden tests or execution results -- iterating against the grader is a stronger intervention and an answer leak.

Drafted by Grok against a written brief, then validated: schema additive and backward-compatible, cost arithmetic checked (2160 calls for 4 models, 5x a raw pass) and grounded in observed fleet-0726b token counts, no model names hardcoded so the set comes from post-fix failure modes rather than the stale pre-fix ones.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
… heat

Three fixes from a day of obliterating Kevin's desktop.

1. LM Studio was JIT-loading at ITS defaults -- 24000 context x 4 parallel slots -- while Gauntlet issues one request at a time at 8192. That sized the KV cache for ~96k tokens to serve 8k, roughly 12x the memory needed; it overflowed VRAM into system RAM and brought the machine to a crawl. It also made the scorecard lie, recording context 8192 for cells that really ran at 24000. Models are now loaded explicitly with --context-length and --parallel 1, plus a TTL backstop so a hard-killed run cannot strand one resident.

2. A mid-stream disconnect killed a 51-cell run outright. client.chat caught ConnectError and HTTPStatusError but not RemoteProtocolError, and merely opening the LM Studio UI is enough to drop a stream. Transport failures are cell outcomes per the error taxonomy, not a reason to abandon hours of finished work.

3. The overlay polled lms ps every 2s, and two copies were running -- a process spawn about once a second, forever. An indicator built to reduce annoyance has no business being a load of its own. Subprocess readings are now cached and refreshed every fifth tick.

It also shows temperature and fan, because that is how Kevin actually notices a run -- he hears it long before he reads a log. Thresholds are calibrated to this machine rather than to a datacentre: idle is ~32% fan and 45-50C, and he heard the fans more on 2026-07-26 than in the preceding year, so 65C/45% is the line. A generic 75C/60% would have read green through the entire day he was complaining, which is worse than no threshold at all.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Long benchmarks belong in the 2am-6am window Kevin nominated rather than being started opportunistically during the day. Resumes by run id, runs without the overlay since nobody is at the machine, and hard-stops at 06:00 -- finishing matters less than the box being silent before he is back at it.

Refuses to start when a run is already going. It exists as a fallback for a run started earlier, so it must never become a second writer to the same run directory; two processes appending to cells.jsonl would corrupt it.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
…ut as UTF-8

The 2026-07-26 fleet run finished with 223 errored cases and twelve unscored cells across its last two models. Root cause: opening the LM Studio desktop app stops its headless HTTP server, and it never came back. lms load kept working -- that talks to the app, not the server -- so VRAM filled, the indicator showed a model resident, and every request was refused. A whole battery burned through in seconds and landed as cells that looked measured.

A ping before the first call to a target now catches it. The target is then recorded as unscored with the reason and the fix (lms server start) rather than being hammered once per case. The run still continues: existing tests correctly caught a first attempt that raised instead, which would have violated the never-abort invariant in CLAUDE.md.

Separately, every lms subprocess call used text=True, which decodes with the locale codec -- cp1252 here -- while lms writes UTF-8 progress output. That killed subprocess's reader thread mid-read, so the command appeared to run and returned nothing usable. Decoded explicitly now, with errors=replace, because an undecodable byte in a progress spinner has no business taking down a benchmark.

Also stops a narrowed resume reporting 51/14 (364%): done holds every cell in the run directory, so it has to be intersected with the current plan.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Three lamps were stacked on screen (six processes -- some sat exactly on top of each other). Two causes. release_overlay_lock deleted the lock file unconditionally, so an orphaned overlay exiting would clear a lock held by the current one and the next run would start another. And _stop_overlay only ran on the happy path, so every hard-killed run today left its lamp behind. The lock is now only dropped by its owner, run() closes the lamp on crash and Ctrl-C too, and the overlay itself refuses to start when a live peer exists -- a borderless window has no title bar, so a stacked one is invisible and you just read whichever is on top.

It also read 'idle -- 1 model loaded' next to '2.7/32GB', because lms ps lists models held by a linked instance on another box. Warning about someone else's GPU is a false alarm here, so the lamp now counts local models only and correctly shows the card free.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
The first real routing report ranked gemma-4-12b at 0.98 above tulu at 0.66. But gemma was scored on 84 of 108 cases and tulu on all 108, and truncation rises with difficulty: gemma lost 7% of T2 and 35% of T4. Its T4=1.00 is a mean taken after discarding the third of T4 it thought hardest about, and its flat tier profile is survivorship, not capability.

The truncated->unscored rule is still right -- a cut-off reply says nothing about ability -- but it makes a biased estimator whenever truncation correlates with difficulty, and the bias runs upward. The leaderboard now carries a  column, flags anything under 90%, and says plainly that a low-coverage row is a ceiling rather than a measurement.

Coverage counts only outcomes that yield no score (truncated, integrity_violation, harness_error). A wrong answer is a measurement and must not shrink the denominator.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Ryfter and others added 10 commits August 21, 2026 02:30
Parse DEVICE from lms ps header-aware output so linked remote hosts
(ITSCM-*) are shown separately and never unloaded via local helpers.

Co-authored-by: Cursor <cursoragent@cursor.com>
Co-authored-by: Cursor <cursoragent@cursor.com>
Co-authored-by: Cursor <cursoragent@cursor.com>
…ting Kevin sign-off)

Co-authored-by: Cursor <cursoragent@cursor.com>
Documents scoring contract for Kevin sign-off; 10 ToC scheduler tests green.

Co-authored-by: Cursor <cursoragent@cursor.com>
Covers multi-critical strongest-first routing, tight-box cheap packing,
and weight-based heavy-before-light scheduling from Kevin's ToC coding.

Co-authored-by: Cursor <cursoragent@cursor.com>
Co-authored-by: Cursor <cursoragent@cursor.com>
Covers empty plans, cost/weight precedence, 5.0 ceiling routing, parallel
tight packing, critical+affinity pinning, heavy→broad VRAM, and defer paths.

Co-authored-by: Cursor <cursoragent@cursor.com>
Documents why boundary tests were added and which scheduling invariants
they protect.

Co-authored-by: Cursor <cursoragent@cursor.com>
Kevin approved ch-c5d002f8232e — unlocks #3–#5 beyond pure algorithms.

Co-authored-by: Cursor <cursoragent@cursor.com>
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants