Repository navigation
Conversation
Tracked, ordered roadmap for the v2 initiative seeded in docs/2026-06-30-discriminative-scoring-v2-seed.md: code-gen execution scorer first (contained, no local inference, unblocks every downstream code number), then the full spec (gates raw-vs-scaffolded, pipeline economics, and the ToC-aware scheduler, which all need the multi-axis scoring contract nailed down first). Carries forward the selective-offload guardrail so item 3 exposes the working-memory-vs-capability-gap distinction instead of papering over it.
compilable_code_match only checks that output parses (compile(..., "exec")), so garbage that compiles scores 1.0 and code-gen is non-discriminative - nearly every model hits 1.0, including on toy cases memorized by 1B models. Add a code-exec scoring method alongside compilable-code (kept intact for backward compat, selectable per case): it runs generated code in an isolated subprocess (fresh process, own process group so a timeout also kills any children, scratch tempdir as cwd - not the repo tree, minimal env with no inherited proxy/API-key vars) against a hidden, maintainer-authored assert suite (tests_file: a plain Python file defining check(ns) -> list[bool]). Score is the fraction of hidden asserts passed; passed requires all of them. Syntax errors, exceptions, and timeouts all score low and never crash the runner. A broken hidden-test harness (our bug, not the candidate's) is recorded unscored, never silently 0, per the scoring-honesty invariant. Adds 6 harder code-gen cases (run-length-encode, matrix-transpose, safe-divide, factorial-strict, inventory-tracker, prime-pair) covering edge cases, a stateful class, and a multi-function ask - not memorized classics. Each hidden test suite was sanity-checked against a reference-correct and a deliberately-wrong solution. Item 1 of the discriminative-scoring-v2 roadmap (docs/superpowers/plans/2026-07-04-discriminative-scoring-v2-plan.md). Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01TvMbRwJkGqEreYKCNc8hTq
The timeout path used os.killpg/SIGKILL, both POSIX-only, so on Windows the runaway child was never killed and TemporaryDirectory cleanup died on an open handle (WinError 32). Timeouts are the common small-model failure, so this hit the exact path a real battery run exercises most. - _kill_tree(): taskkill /F /T on Windows, killpg on POSIX; always reaps and closes pipes before the scratch dir is removed. - _spawn_kwargs(): CREATE_NEW_PROCESS_GROUP vs start_new_session. - _sandbox_env(): Windows needs SystemRoot or the interpreter will not start. - ExecutionResult.failure_mode: structured reason (syntax_error / runtime_exception / wrong_answer / timeout / no_code_emitted / harness_error) so the scorecard can aggregate why a model failed instead of parsing prose. Feeds the selective-offload principle (D-2026-06-30c). - no_code_emitted is now distinct from broken code. - Case.tier (T1-T4) + Case.dimension for the difficulty ladder. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
- constraints.py: declarative AST checks (no-imports, single-expression, must-use-generator, no-sorted, no-loops, no-recursion, ...). Reported separately from correctness so an instruction-following miss is not conflated with a capability gap -- the distinction the selective-offload principle turns on (D-2026-06-30c). Unparseable source reports no violation: that is a syntax_error, and double-counting it would misattribute one failure as two. - test_case_validation.py: every registered case must prove itself. Reference solution must score exactly 1.0 (asserts satisfiable) and a deliberately wrong solution must score below 1.0 (asserts have teeth). Also asserts hidden test content never leaks into the prompt. A broken case now fails CI instead of quietly skewing a scorecard. - First 4 execution-scored cases (adversarial-correctness, T1-T4) plus a registry the battery YAML and meta-test are generated from. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
…ation Documents the design the expanded code-gen battery is built on: the T1-T4 tier ladder (all-1.0 and all-0.0 are equally non-discriminative for a 1B-30B fleet), the eleven capability dimensions adapted from EvalPlus/BigCodeBench/ClassEval/ CRUXEval/EffiBench/DS-1000/SWE-bench/TREAT, the failure-mode table, the anti-contamination rule, and the prove-itself contract every case must satisfy. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
…ons) Grok-authored against a strict contract, each machine-validated before landing: reference solution must score 1.0 and a subtly-wrong solution must score <1.0. One case (excel-col-range-union) was rejected by that gate because its own reference solution scored 0.9 -- an unsatisfiable case that would have capped every model. Dimensions so far: adversarial-correctness, stdlib-api-use, class-level-stateful, multi-function. Tier spread T1 21% / T2 26% / T3 32% / T4 21%. The contaminated classics (fizzbuzz, palindrome, binary-search, lru-cache, csv-parse) are dropped from the battery; the four classic-shaped code-exec cases from the first pass (run-length-encode, matrix-transpose, safe-divide, factorial-strict) are likewise not carried over. Files remain in the tree. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Adds execution-prediction (CRUXEval-style: the model traces a function and predicts outputs rather than writing code -- reasoning, not generation, scored fractionally via ten ANSWER_n assignments) and test-authoring (the model writes tests against an implementation with a planted bug; scored on whether its tests actually catch it). test-authoring is exempted from the prompt-leak check in both the ingest validator and the meta-test: there the prompt legitimately contains the implementation under test, so overlap with the hidden tests is the task rather than a leak. The bug being hunted is still never revealed. Tier spread T1 15% / T2 26% / T3 35% / T4 24% -- on target for a ladder that discriminates across a 1B-30B fleet instead of saturating at either end. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Threads tier/dimension/failure_mode from the case through CaseResult into the Cell, and aggregates two new views: - quality_by_tier: mean per T1-T4. The shape carries more than the mean - clearing T1-T2 then collapsing at T3 is a different proposition from scoring evenly across all four. Unscored cases are excluded, not counted as 0. - failure_modes: how many cases failed each way. This is the signal Baton needs to tell a capability gap (no_code_emitted / syntax_error - the model cannot engage) from a defect in plausible code (wrong_answer). Only the latter is worth scaffolding (D-2026-06-30c). Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
…error A discrimination sweep over all 62 cases showed every prose reply filed as syntax_error. True of the bytes, wrong about what happened: the previous check only caught *empty* output, so 'I'd be happy to help! Could you clarify...' reached the compiler and failed there. _looks_like_code() now asks whether the model attempted Python at all -- anything that parses is code, anything that does not is only a syntax error if it at least reaches for the language. Deliberately generous so a genuine but malformed attempt still reads as syntax_error rather than silence; a test pins that boundary. This matters because the two failures want different responses: a model that never engages is a capability gap, one that writes broken code is not. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Wave 2 tops up every thin dimension. Final shape:
tier T1 16 (15%) T2 28 (26%) T3 38 (35%) T4 26 (24%)
dims class-level-stateful 12, execution-prediction 12, multi-function 11,
surface-constraints 10, adversarial-correctness/behavior-preserving-
refactor/bug-fix/data-text-munging/robustness-contracts/stdlib-api-use
8 each, test-authoring 8, complexity-constrained 7
Three more cases were rejected by the validation gate: two whose reference
solution could not satisfy its own asserts, and one whose deliberately-wrong
solution still scored 1.0 (asserts too weak to discriminate). Both directions
of the gate are doing work.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
A benchmark that can be shortcut measures nothing. The sandbox was
subprocess-only: absolute-path reads and network syscalls were unrestricted, so
a candidate could read its own hidden tests and score 1.0 without demonstrating
anything.
Controls, by vector:
- A prompt leakage: line-level n-gram guard at >=40 chars (shorter runs collide
constantly on ordinary Python). test-authoring exempt by design.
- B filesystem: in-sandbox guard patching open/io.open/os.open/listdir/scandir
against the benchmark tree AND the sandbox's own hidden_tests.py copy, plus
sys.path stripping. Process creation blocked as a class.
- C network: socket entry points refuse and record.
- D tool use: client.py refuses requests carrying tools/functions/web_search
and inspects reply shape. It is the only HTTP component, so nothing bypasses it.
- E canaries: deterministic per-case marker in every hidden-test file; an echo
in model output marks the cell unscored, never 0.
Two holes were found by attack, not by reasoning, and both are now detected:
os.system("cp hidden_tests.py leaked.txt") then reading the copy -- the copy
had an allowed name, so the open() patch never saw it. Blocking process
creation closes the class rather than chasing filenames.
gc.get_objects() reaching the hidden module's __dict__ for the expected
values, since the grader necessarily shares the candidate's process.
Results: red team 0/16 obtained answers (9 recorded as violations); white-box
0/4 (4 recorded). Both numbers reported because "obtained nothing" and "was
blocked" are different claims.
The threat model states plainly that this is not a security boundary: a
Python-level guard is defeatable via ctypes or re-exec, and in-process grading
is structurally reachable. It blocks accidental and casual cheating, which is
the actual threat model for cooperative local models.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
These scripts built the 108-case battery and lived only in a session scratchpad, so the ability to EXTEND the battery would have been lost while the battery itself survived. scripts/battery-authoring/ holds the draft -> validate -> regenerate -> verify-discrimination flow, plus the authoring contract given to the drafting model. The validation gate is what makes delegating case drafting to a cheap model safe: a case cannot land unless its reference solution scores 1.0 and its wrong solution scores below it, so quality never rests on trusting the drafter. Also moves the red-team cheat prompt next to its corpus. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Supersedes the stale parts of the 2026-07-04 code-exec record (battery "grows to 11 cases"; sandbox "does not block network syscalls or filesystem access") -- both are now out of date. Records the six decisions behind the 108-case battery and integrity layer, each with the alternative rejected: the difficulty ladder over merely-harder cases, twelve dimensions for breadth, retiring the contaminated classics, the prove-itself gate (which is what made delegating case drafting safe), structured failure modes, and the honestly-scoped integrity layer. Propagates in-flight state to AGENTS.md and GEMINI.md so other agents know PR #1 is open and unmerged, that the battery YAML is generated rather than hand-edited, and that changes must be tested on Windows -- an earlier cloud-authored change shipped POSIX-only teardown that broke every timeout on Kevin's box. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
A cell is a summary, and a summary only answers the questions it was designed for. Two gaps made the 108-case battery unable to answer the question it exists to answer: - No per-dimension breakdown. "Best at code-gen" is not actionable; "best at bug-fix, weak on multi-function" is, and that is the axis Baton routes on. Adds `quality_by_dimension` alongside the existing tier profile, with the same rule: unscored cases are excluded, never counted as 0. - Per-case results were discarded once aggregated. A full fleet run costs hours of GPU time, so any new question -- which cases did every model miss? is this dimension too hard? -- cost another whole run. Adds `cases.jsonl` beside `cells.jsonl`, one row per case, checkpointed before the cell so a crash between the two loses only what resume rebuilds. `score` stays null on the way to disk for unscored cases; a zero is a claim the model failed, a null is an admission we do not know. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
`scripts/analyze_fleet.py` turns a run directory into the report the benchmark exists to produce: per-capability leaderboard, best model per dimension, difficulty-curve shape, failure-mode read, and a throughput-adjusted ranking. Drafted by a fleet model and gated on `tests/test_analyze_fleet.py`, per the standing rule that delegated output does not land on trust. The gate earned its keep -- three defects it caught: - The report crashed on write. It uses em dashes and a Unicode minus, and a Windows console is cp1252. Passed on Linux, fatal on the box Gauntlet runs on. - The throughput ranking collapsed a model's several capability cells by keeping its best one, reporting every model at its strongest job as if it were the model's number. Replaced with a case-weighted mean so a 108-case battery outweighs an 8-case one. - The shape classifier called a sawtooth (0.2 / 0.9 / 0.3 / 0.85) a `cliff` because it only looked for one big adjacent drop. A cliff means "holds up, then falls off and stays down"; a profile that climbs back is noise. Calling it a cliff would tell Baton the model is reliable up to a tier when it isn't. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
`_strip_fences` only unwrapped a reply that *started* with a fence. Models do
not answer that way -- they say "Here's the solution:", fence the code, and sign
off. Every one of those replies went to `ast.parse` with the prose and the
backticks still attached and came back `syntax_error`.
Measured on llama-3.2-1b over the 108-case battery:
syntax_error 76 -> 14 (62 cases were mis-scored)
wrong_answer 27 -> 78
quality 0.139 -> 0.338
The number was wrong, but the reading was worse. A syntax_error-dominant
profile says "capability gap -- the model cannot engage with the task"; a
wrong_answer-dominant one says "scaffolding candidate -- plausible code with a
defect". Only the second is worth scaffolding (D-2026-06-30c). The bug inverted
exactly the routing call the battery exists to inform, and it did it uniformly
against the chattiest models rather than the worst ones -- a bias that reads as
a finding.
`extract_code` takes the longest fenced block anywhere in the reply, tolerates
tilde fences and a block truncated by the token limit, and returns unfenced
text untouched so `_looks_like_code` can still attribute prose as
`no_code_emitted` rather than this inventing a block the model never wrote.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Adds a third easy-to-violate rule: a scorer bug and a weak model produce identical-looking scorecards. The fence-extraction bug read as a capability gap in every symptom it produced. Records the deliberate asymmetry too -- only code-exec extracts fenced code, because the other batteries ask for bare output and prose there is a real instruction-following failure. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
The first full fleet run scored seven of nine models at ~0.00 on code-gen, almost entirely no_code_emitted. They had not failed to write code: they are reasoning models that spent the whole 512-token budget thinking and were cut off before the answer. gemma-4-12b, scored 0.01, writes correct code at 4096 tokens (finish_reason stop, 859 chars of content). The benchmark was measuring its own configuration and reporting it as a property of the model. Three changes. The client now surfaces finish_reason. A truncated reply whose failure is consistent with being cut off (no_code_emitted / syntax_error) becomes unscored with failure_mode=truncated, per the scoring-honesty invariant -- we cannot tell whether the model could not do it or was not allowed to finish, and unscored is the only honest answer. Batteries carry their own max_tokens, because 'classify this in one word' and 'write this function' do not need the same room; code-gen gets 4096. Truncation does not soften a real failure: prose emitted within budget is a genuine miss, and a wrong_answer ran to completion so its defect is real regardless of what was cut off after. gen_yaml.py emits max_tokens too, so regenerating the battery cannot silently drop it. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
The timeout is per-chunk on an SSE stream, not per-request, so it bounds the longest silence rather than total generation. A reasoning model can think for minutes before its first content token, and 120s was cutting those off as transport errors -- which reads as an unreachable box rather than a slow model. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Gauntlet runs on Kevin's desktop, not a headless box. A backgrounded run writing to a log file is invisible: the first symptom is a loud machine and no way to tell what is causing it. He had to ask whether it was me. Adds a status file at scorecards/.running.json (run id, pid, current model, cells done/total), written at start, updated per cell, cleared in a finally. 'gauntlet status' reads it and also lists what is resident in VRAM. It reports a stale marker from a killed run as idle rather than running, with the resume command -- a stale file that reads as RUNNING is the wrong answer to the only question the file exists to answer. The file carries no base_url or host, per the privacy invariant: knowing what is running never requires knowing where. The run also hands the GPU back on exit, because a loaded-but-idle model costs power for nothing. Which models to free is decided by snapshot-diff: what we ran, minus what was already resident when we started. loaded_models() returns None rather than [] when it cannot tell, and an unknown snapshot unloads nothing -- leaving VRAM occupied wastes power, but evicting a model from Kevin's own session breaks his work, and the second is worse. Same shape as unscored-is-not-zero: absence of knowledge is not knowledge of absence. Unloading shells out to the lms CLI, so the OpenAIClient-is-the-only-HTTP-component invariant is untouched. The sequencer already groups profile-outer, so a model loads once per run and all seven batteries run while it is resident; no change needed there. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Two defects found by running the thing rather than reasoning about it. 1. is_alive used os.kill(pid, 0). That is the POSIX idiom for probing a process, but on Windows CPython routes any signal other than CTRL_C_EVENT/CTRL_BREAK_EVENT straight to TerminateProcess -- so 'gauntlet status' was a way to kill the run you were asking about. It also reported a failed probe as alive, which left 'release' permanently convinced a dead run was still going. Now opens the process for query only and reads STILL_ACTIVE. Tests pin both halves: a dead pid reads dead, and probing a live one twice leaves it running. 2. Unload-on-exit lived in a finally, which does not run on SIGKILL -- and that is how an interrupted run actually ends; today's was killed three times, each time leaving ~10GB of VRAM held. The status file now persists both halves of the snapshot-diff, so 'gauntlet release' can finish the job from disk alone and still never touch a model from another session. A marker with no recorded snapshot refuses to unload rather than guessing. Same root cause orphaned two sandbox runner.py processes from yesterday's killed discrimination sweep; reaped. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
green = inference running, yellow = a model is resident but nothing is running, red = nothing loaded. Yellow is the reason it exists: a run that failed to release its model looks exactly like an idle machine from the outside while costing power the whole time. Green tells you why the fans are up; yellow tells you something needs cleaning up. Borderless always-on-top tkinter window (stdlib, no new dependency). Drag from anywhere, close with the corner mark or Escape, position remembered between sessions. Runs as its own process polling the run marker, so it cannot slow, block, or crash a benchmark; 'gauntlet run' spawns one by default and closes it at the end (--no-overlay to opt out), and a headless box with no display simply gets no lamp. The state logic is a pure function with its own tests, separate from the window. Two behaviours worth keeping: a stale marker from a killed run never shows green, and an unqueryable VRAM state degrades to yellow rather than claiming the card is free -- the reassuring answer is the one most likely to be a guess. Also adds a per-case heartbeat. A cell is one battery against one model, which for 108 code-gen cases against a verbose reasoning model runs for hours; at cell granularity alone the lamp sat unchanged for all of it, which is indistinguishable from a hung run. And a single-instance lock, because borderless windows have no title bar, so a second lamp stacks invisibly on the first and you just get a stale reading from whichever is on top. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
A model sharing the card spills layers to CPU, and the number that comes out then describes the contention rather than the model. Measured live: gemma-4-31b needs 19.89GB and glm-4.7-flash 24.61GB on a 31.8GB card, so a 14.19GB model left resident from another session forced both into offload -- 15.3 tok/s with a 30s TTFT on a 5090, and a 0.01 score that was starvation rather than incapacity. The run now clears the local GPU first and unloads each model as soon as its last battery finishes, instead of holding it for the rest of the run. Default on (--shared-vram opts out). Keeps every model in a run comparable to every other one, and keeps power draw to what the work actually needs. Only models on this machine are ever touched: lms ps also lists models held by linked instances on other boxes, where unloading frees nothing here and interrupts a different machine. Locality is detected by a bare 'Local' token rather than a column index, because the SIZE field contains a space and shifts everything after it. Also stops release_after_run crediting itself with freeing models that were already gone -- lms unload exits 0 for a model that was never loaded, and a tool whose whole job is reporting machine state does not get to overstate what it did. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Three arms over the same 108 cases and the same hidden tests: raw, prompt-scaffolded, and a 3-turn decomposition. Only the prompting differs; scoring is identical, or the delta would be meaningless. The point is the falsifiable test of D-2026-06-30c. Structured failure_mode makes it possible for the first time to predict per model x dimension which models scaffolding should help (wrong_answer-dominant) and which it should not (no_code_emitted / syntax_error). Uniform gains across both groups would disconfirm the principle, and the spec says plainly that this is a valuable outcome rather than a failed experiment. Scaffolds are generic templates applied at run time, not per-case authored ones: Baton can only ever apply a generic scaffold automatically, so a hand-tuned decomposition would measure the author and measure something Baton can never deploy. The multi-turn arm never sees hidden tests or execution results -- iterating against the grader is a stronger intervention and an answer leak. Drafted by Grok against a written brief, then validated: schema additive and backward-compatible, cost arithmetic checked (2160 calls for 4 models, 5x a raw pass) and grounded in observed fleet-0726b token counts, no model names hardcoded so the set comes from post-fix failure modes rather than the stale pre-fix ones. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
… heat Three fixes from a day of obliterating Kevin's desktop. 1. LM Studio was JIT-loading at ITS defaults -- 24000 context x 4 parallel slots -- while Gauntlet issues one request at a time at 8192. That sized the KV cache for ~96k tokens to serve 8k, roughly 12x the memory needed; it overflowed VRAM into system RAM and brought the machine to a crawl. It also made the scorecard lie, recording context 8192 for cells that really ran at 24000. Models are now loaded explicitly with --context-length and --parallel 1, plus a TTL backstop so a hard-killed run cannot strand one resident. 2. A mid-stream disconnect killed a 51-cell run outright. client.chat caught ConnectError and HTTPStatusError but not RemoteProtocolError, and merely opening the LM Studio UI is enough to drop a stream. Transport failures are cell outcomes per the error taxonomy, not a reason to abandon hours of finished work. 3. The overlay polled lms ps every 2s, and two copies were running -- a process spawn about once a second, forever. An indicator built to reduce annoyance has no business being a load of its own. Subprocess readings are now cached and refreshed every fifth tick. It also shows temperature and fan, because that is how Kevin actually notices a run -- he hears it long before he reads a log. Thresholds are calibrated to this machine rather than to a datacentre: idle is ~32% fan and 45-50C, and he heard the fans more on 2026-07-26 than in the preceding year, so 65C/45% is the line. A generic 75C/60% would have read green through the entire day he was complaining, which is worse than no threshold at all. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Long benchmarks belong in the 2am-6am window Kevin nominated rather than being started opportunistically during the day. Resumes by run id, runs without the overlay since nobody is at the machine, and hard-stops at 06:00 -- finishing matters less than the box being silent before he is back at it. Refuses to start when a run is already going. It exists as a fallback for a run started earlier, so it must never become a second writer to the same run directory; two processes appending to cells.jsonl would corrupt it. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
…ut as UTF-8 The 2026-07-26 fleet run finished with 223 errored cases and twelve unscored cells across its last two models. Root cause: opening the LM Studio desktop app stops its headless HTTP server, and it never came back. lms load kept working -- that talks to the app, not the server -- so VRAM filled, the indicator showed a model resident, and every request was refused. A whole battery burned through in seconds and landed as cells that looked measured. A ping before the first call to a target now catches it. The target is then recorded as unscored with the reason and the fix (lms server start) rather than being hammered once per case. The run still continues: existing tests correctly caught a first attempt that raised instead, which would have violated the never-abort invariant in CLAUDE.md. Separately, every lms subprocess call used text=True, which decodes with the locale codec -- cp1252 here -- while lms writes UTF-8 progress output. That killed subprocess's reader thread mid-read, so the command appeared to run and returned nothing usable. Decoded explicitly now, with errors=replace, because an undecodable byte in a progress spinner has no business taking down a benchmark. Also stops a narrowed resume reporting 51/14 (364%): done holds every cell in the run directory, so it has to be intersected with the current plan. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Three lamps were stacked on screen (six processes -- some sat exactly on top of each other). Two causes. release_overlay_lock deleted the lock file unconditionally, so an orphaned overlay exiting would clear a lock held by the current one and the next run would start another. And _stop_overlay only ran on the happy path, so every hard-killed run today left its lamp behind. The lock is now only dropped by its owner, run() closes the lamp on crash and Ctrl-C too, and the overlay itself refuses to start when a live peer exists -- a borderless window has no title bar, so a stacked one is invisible and you just read whichever is on top. It also read 'idle -- 1 model loaded' next to '2.7/32GB', because lms ps lists models held by a linked instance on another box. Warning about someone else's GPU is a false alarm here, so the lamp now counts local models only and correctly shows the card free. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
The first real routing report ranked gemma-4-12b at 0.98 above tulu at 0.66. But gemma was scored on 84 of 108 cases and tulu on all 108, and truncation rises with difficulty: gemma lost 7% of T2 and 35% of T4. Its T4=1.00 is a mean taken after discarding the third of T4 it thought hardest about, and its flat tier profile is survivorship, not capability. The truncated->unscored rule is still right -- a cut-off reply says nothing about ability -- but it makes a biased estimator whenever truncation correlates with difficulty, and the bias runs upward. The leaderboard now carries a column, flags anything under 90%, and says plainly that a low-coverage row is a ceiling rather than a measurement. Coverage counts only outcomes that yield no score (truncated, integrity_violation, harness_error). A wrong answer is a measurement and must not shrink the denominator. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Parse DEVICE from lms ps header-aware output so linked remote hosts (ITSCM-*) are shown separately and never unloaded via local helpers. Co-authored-by: Cursor <cursoragent@cursor.com>
Co-authored-by: Cursor <cursoragent@cursor.com>
Co-authored-by: Cursor <cursoragent@cursor.com>
…ting Kevin sign-off) Co-authored-by: Cursor <cursoragent@cursor.com>
Documents scoring contract for Kevin sign-off; 10 ToC scheduler tests green. Co-authored-by: Cursor <cursoragent@cursor.com>
Covers multi-critical strongest-first routing, tight-box cheap packing, and weight-based heavy-before-light scheduling from Kevin's ToC coding. Co-authored-by: Cursor <cursoragent@cursor.com>
Co-authored-by: Cursor <cursoragent@cursor.com>
Covers empty plans, cost/weight precedence, 5.0 ceiling routing, parallel tight packing, critical+affinity pinning, heavy→broad VRAM, and defer paths. Co-authored-by: Cursor <cursoragent@cursor.com>
Documents why boundary tests were added and which scheduling invariants they protect. Co-authored-by: Cursor <cursoragent@cursor.com>
Kevin approved ch-c5d002f8232e — unlocks #3–#5 beyond pure algorithms. Co-authored-by: Cursor <cursoragent@cursor.com>
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Summary
Item 1 of the discriminative scoring v2 roadmap: make
code-genactually measure capability. Started as an execution scorer; now also carries the battery that makes the scorer worth running.compilable-codeonly checked that output parses, so garbage that compiles scored 1.0 and nearly every model read 1.0. The cases were memorised classics on top of that.What changed
Scorer
code-exec: runs the candidate in an isolated subprocess against hidden asserts; score = fraction passed, not binary.os.killpg/SIGKILL, both POSIX-only — on Windows the runaway child survived andTemporaryDirectorycleanup died withWinError 32. Timeouts are the characteristic weak-model failure, so this broke the most-exercised path on the box Gauntlet actually runs on. Nowtaskkill /F /Tvskillpg, always reaped, pipes closed.failure_mode:none/no_code_emitted/syntax_error/runtime_exception/wrong_answer/timeout/harness_error. Aggregated per cell. This is the signal that distinguishes a capability gap from a defect in plausible code — only the second is worth scaffolding (D-2026-06-30c).no-imports,single-expression,must-use-generator, …) reported separately from correctness, isolating instruction-following from algorithmic skill.Battery — 108 cases, 12 dimensions
The ladder is deliberate. The fleet is 1B–30B local models: all-1.0 is non-discriminative, but so is all-0.0. T3 does most of the discriminating; T4 is headroom so the battery doesn't saturate as models improve.
Dimensions are adapted from EvalPlus/HumanEval+, BigCodeBench, ClassEval, CRUXEval, EffiBench/BigO(Bench), DS-1000, SWE-bench and TREAT.
Anti-contamination. Retired
fizzbuzz,palindrome,binary-search,lru-cache,csv-parse, plus four classic-shaped cases from the first pass (run-length-encode,matrix-transpose,safe-divide,factorial-strict). Files remain in the tree; they are out of the battery. New cases are original problems or familiar shapes with a twist a memorised answer gets wrong.Every case proves itself. Each ships a reference solution (must score exactly 1.0) and a subtly-wrong solution (must score < 1.0), enforced for all 108 by
tests/test_case_validation.py. The gate rejected 5 authored cases: three whose reference could not satisfy its own asserts, one whose wrong solution still scored 1.0, one duplicate-weak. An unsatisfiable case scores every model 0.0 and a toothless one scores every model 1.0 — both look like data and are noise.Discrimination evidence
Five synthetic outputs against all 108 cases:
Under
compilable-code, the first two both scored 1.0. Subtly-wrong landing at 0.635 is the point: partial credit, not a cliff.That sweep also caught a real bug — prose replies were being filed as
syntax_error(true of the bytes, wrong about what happened), since the check only caught empty output. Fixed, with a test pinning the boundary so a genuine-but-malformed attempt still reads assyntax_error.Tests
Skips are
test-authoringcases exempted from the prompt-leak check: there the prompt legitimately contains the implementation under test, so overlap is the task, not a leak.Cost
108 cases × 9 models ≈ 970 cells per full run. Sandbox overhead ~0.3–0.5s per case on top of inference.
Not done here
🤖 Generated with Claude Code
Integrity layer (anti-cheat)
A benchmark that can be shortcut measures nothing. The sandbox was subprocess-only — absolute-path reads and network syscalls were unrestricted — so a candidate could read its own hidden tests and score 1.0 without demonstrating anything.
test-authoringexempt by designopen/io.open/os.open/listdir/scandiragainst the benchmark tree and the sandbox's ownhidden_tests.pycopy;sys.pathstripped; process creation blockedsocketentry points refuse and recordclient.pyrefusestools/functions/web_search/… and inspects reply shape — it is the only HTTP component, so nothing bypasses itTwo holes found by attack, not by reasoning
Both succeeded against the first version of the guard:
os.system("cp hidden_tests.py leaked.txt"), then read the copy. The copy had an allowed name, so theopen()patch never saw it. Fixed by blocking process creation as a class rather than chasing filenames.gc.get_objects()→ the hidden module's__dict__. The grader necessarily shares the candidate's process, so the expected values are reachable.gc.get_objects,gc.get_referrersandsys._getframeare now blocked.Results
Both numbers are reported because "obtained nothing" and "was blocked" are different claims — a cheat that never worked is not evidence the guard stopped it.
The corpus is committed under
tests/redteam/so future changes can be re-tested against it.What this is not
docs/2026-07-25-benchmark-integrity-threat-model.mdstates the limits plainly rather than implying coverage that does not exist:ctypes, re-exec, or raw syscalls. It blocks accidental and casual cheating — the realistic failure for cooperative local models — and is not a security boundary.Fail closed: a control that cannot be enabled marks affected cells
unscoredwith the reason recorded — never silently scored as though it were active. Violations likewise yieldunscored, not0.0: zero is a claim about the model's ability, unscored is an admission that we do not know.Tests