feat(bundle): the campaign-record bundle — one conformance set, one presence question, and a record set that has to be checkable - #870
Conversation
|
Important Review skippedToo many files! This PR contains 193 files, which is 93 over the limit of 100. To get a review, reduce the PR to 100 files or fewer by splitting it into smaller PRs or changing its base branch. Upgrade to a paid plan to raise the limit. This review couldn't start because sufficient usage credits or metered capacity aren't available. Add credits or update usage-based reviews in the billing tab, then retry. ⚙️ Run configurationConfiguration used: Path: .coderabbit.yaml Review profile: CHILL Plan: Team Run ID: ⛔ Files ignored due to path filters (1)
📒 Files selected for processing (193)
You can disable this status message by setting the Thanks for using CodeRabbit! It's free for OSS, and your support helps us grow. If you like it, consider giving us a shout-out. Comment |
66fd221 to
5b06ff5
Compare
CLOUD-1142 landed the definition — `ready::Grammar::key_of` and `keys_in` own the three axes. What it could not land is the obligation: the nineteen governed shell consumers still re-derive the key, each retires whole under CLOUD-1164, and nothing made its successor's tier import the examples rather than re-type them. `crates/batten/tests/it/issue_key.rs` exports `conformance()` — the row's fixed seven, decided by the engine's own reader resolved from the COMMITTED `[[pattern]]` table, so a registry row's drift reddens the integration tier and not only the crate's own unit tests. `no-tracker-key-in-core` carried ONE spelling of the key expression, which refuses the way the twenty measured derivations happened to be written and waves through every other way of writing the same thing. It is an alternation now, and `no-tracker-key-in-modules` is its twin over `policy/**` — a `forbid` rather than a ratchet there, because that tree carries zero derivations to wait out. Both are asserted by the committed-ruleset census, whose full-stdout equality is what tells a mis-globbed row from a working one. The declared mutation is bound in the test file, where `obligations-bound` reads it, and declared again beside the reader in `ready.rs` — `engine-ready` joins `$MUTANT_GATES`, so `mise run mutant` APPLIES it rather than only binding it, closing the gap `commit_arm_sequencing.rs` records for its own row. Refs: CLOUD-761, CLOUD-1142, CLOUD-1164, CLOUD-418 Admits: 1f335ff6f89a7f7f13207996572b4913ddef80e9e0f7cc9629b0ff46f80531e1 Admits-rule: protected-mutation Admits-verdict: path write refused Admits-subject: batten.toml Admits-anchor: call:a16725c3350d90c01673dd1a521cb1e1f3e4057d Admits-epoch: 55191757a873a6121f5968407711eb9d18456d12f8af05fd7850814f3823f9f6 Admits-author: alec@wenzowski.com Admits-prev: b5b2f6da78d6ac648923198e0d3f1cc0283fe4ce68f2d4c5276f934fa35e445a Admits-answer-lost: CLOUD-761's deliverable. `no-tracker-key-in-core` refuses exactly one spelling, the literal `CLOUD-[0-9]`, and `policy/**` carries no such gate at all — so a module or crate file composing `CLOUD-\d` is a second authority over the one definition of an issue key and nothing refuses it. A class a second spelling escapes is not a class; twenty derivations accumulated because nothing refused the twenty-first before it was typed. Admits-answer-precondition: A `[[rule]]` row is declared in batten.toml and nowhere else, so no other surface can express this change. It adds one `forbid` row over `policy/**` (`no-tracker-key-in-modules`) and widens `no-tracker-key-in-core`'s regex from one spelling of the key expression to three; a reviewer reads both directly in the diff as one new row and one changed line. Admits-answer-rejected-route: config read first does not apply: the file has been read, and the finding IS that the committed row's regex covers one spelling of three. patch run first does not apply either — there is no patch surface for a `[[rule]]` row, the row IS the config.
… locked ones `lock-complete` asks two questions and one exemption was wired into both. Whether an entry LOCKS a url is genuinely inapplicable to a backend that resolves through its own package manager — that is what `locks_nothing` is for. Whether a tool HAS an entry applies to every backend without exception, because `mise install --locked` demands a row whatever the row contains. `lock-tool-missing` inherited the exemption anyway, and its own comment called that "a LENIENCE here rather than a necessity". Measured on CLOUD-580's branch, in a backend the clause exempted and one issue after CLOUD-333 closed: `"pipx:ntia-conformance-checker" = "5.0.3"` with no row passed at exit 0 and `verify` green twice on two SHAs, then took every mise-action job red at the install step — `zizmor` and `commit-lint` included, since `--locked` validates the whole file rather than the job's install list. The edit is one conjunct. `lock-tool-unlocked`'s url arm keeps the identical call, and it is a genuinely different call: it reads the backend off a row that EXISTS, where this clause has no row to read. `key_backend` goes with it rather than being left dead — it existed only to feed the exemption. ## The fixture that had been hiding behind the exemption One load-time test needed changing, which the row named as the signal that a removal reached the wrong clause. It did not, and the reason is worth the sentence: `test_a_backend_that_cannot_lock_is_exempt_from_locking_nothing` keyed its lock entry `core:rust` while its manifest said `rust`, a pairing mise does not produce — the committed `mise.lock` says `[[tools.rust]]` with `backend = "core:rust"`. The presence lookup is by the lock's own key, so the mismatch was always there and the exemption short-circuited before anything could reach it. The fixture now spells the entry the way the tool writes it and tests the url question alone. Verified against the real tree as well as the fixtures: `batten check --rule lock-complete` is silent over this repository, so the narrowing changed no verdict on a correct tree — every exempt-backend pin here already carries a row. The declared mutation re-creates the removed blind spot by re-exempting one prefix on the presence clause; applied, it reddens only the new case and leaves the seventeen others green. Refs: CLOUD-611, CLOUD-333, CLOUD-580, CLOUD-418
…cision log
CLOUD-268 moved model, harness and session identifiers out of durable public
commit text, and `batten attribution` enforces that suppression over produced
commits. Suppression alone trades an unreliable public signal for nothing
unless the operational question stays answerable, so this is the read side:
`decision::attribution_for(repo_root, sha)`.
The join key and the payload were both already there — `Anchor.commit`, whose
own doc names this row, and `Caller`'s three fields, reserved by CLOUD-133.
What was missing is the query.
## A grouping, never a merge
A commit can carry many records, and they can disagree: two sessions, or a tree
that went dirty between two calls. Folding them into one row would invent a
fact no record carries, so each distinct `(caller, dirty)` pair is its own row
with the count that carried it. `dirty` is part of the key rather than a column
because it changes what the row claims — a record taken against a dirty tree
describes a call the commit does not fully contain.
## Absent is not unknown
No rows means the log holds nothing anchored to this sha: the could-not-look
case, and the ordinary state of a commit made outside a mediated session. A row
valued `unknown` means a record exists and the host exposed no identity.
Collapsing them makes a gap read as a fact — CLOUD-251's trap in this surface —
so both halves are asserted in one case, where neither can drift from the other.
The rows carry each record's OWN `Caller` rather than one rebuilt from its
tokens. `Provenance::from_host` degrades empty and whitespace; it is
`Provenance`'s `Deserialize` that maps the `unknown` token back to `Unknown`. A
reconstruction through `from_host` would turn a degraded field into
`Declared("unknown")` — the query and the record disagreeing about what a
degraded field is, which is the one distinction this row exists to keep. Caught
while writing it, not by a test.
## Rule 4 structurally
`subject` and `context` are not fields of the answer, so no rendering choice can
leak them. The pointer-only case plants a real path and real context bytes in
the record and asserts none of it reaches the rendered answer, so it would fail
if a later author added such a field "for context".
## The declared mutation, and the one that survived first
`degraded-field-not-unknown` stops rendering the degraded token, so an
undeclared field serializes empty and reads back as `Declared("")`. Applied, it
reddens the degraded case alone and leaves the other three green.
Its first spelling mutated `Provenance::from_host` and SURVIVED:
`Caller::undeclared()` builds `Provenance::Unknown` directly and never calls
`from_host`, so the case never walked the mutated line. The file header records
that, because choosing a mutation over a path the case does not take is exactly
how a declaration becomes coverage that cannot fail.
Refs: CLOUD-275, CLOUD-133, CLOUD-268, CLOUD-251, CLOUD-418
…issing A `[[rule]]` row is a classifier with two failure modes and the corpus observed one. `expected.in` pins the whole of stdout, so a rule that stops firing does change those bytes — but it fails as a stdout diff, which names the fixture and not the rule, and it says nothing at all about a rule firing on a line nobody meant it to. The four instances in `batten.toml`'s own history are all that untested half: a comment narrating a banned command, literals dropped because they fired on ordinary prose, a leading quote carried purely to suppress a false positive, and a self-match hazard nothing pinned. Measured (CLOUD-310): of 40 lines a literal row reported across `mise-tasks/`, 8 were comments — 20% noise the suite could not see. ## The expectation lives in the case `# batten-case: <rule-id> clean|violating` declares the line that FOLLOWS it — `#MUTANT`'s shape, for its reason: one authority, adjacent to its subject, with no second file to keep in agreement. The three candidates the row listed each fail on a rule this repository already holds, and the Ready block chose this fourth. Four outcomes, two of them named failures: `reported`, `validated`, `noisy`, `missing`. The runner is `fixture_repos.rs` extended rather than a new program — a new `mise-tasks/` file is born governed, so its own first bug fix would have had nowhere to land. Output is the case's `path:line rule-id` and the outcome token, never the matched bytes: a false-positive report whose subject is a matched literal is exactly where a checker leaks what it was scanning. ## Shown able to fail, in the direction the row is about `forbid-quote-load-bearing` reproduces `no-source-built-tool`'s shape without depending on this repository's own rule table. Dropping the leading quote from its pattern makes the rule report the comment explaining itself, and the suite now fails as `pins.toml:3 no-built-backend noisy` rather than as a stdout diff. The quote is pinned as load-bearing instead of remembered in a comment. `forbid-exclude-comment`'s two comment lines were the historical false positive and were clean only by someone's memory; they are declared cases now. The scorer is negatively self-tested in both directions, and a finding one line off scores as `missing` rather than as a pass. ## The self-match convention gains a test Rule ids deliberately do not name their own literals, because a finding's pointer output lands in a fixture inside the rule's own glob. That was a convention with nothing behind it, and a marker naming a rule id is exactly what a reintroduced self-match would break — so it is asserted over the corpus now. ## One defect this change caused and fixed The scratch path `fixture-repos/<name>` was unambiguous with one consumer and I added two more. Measured: the scoring test passed alone and scored `missing` beside its sibling, reading a tree the other was still writing. Process-per-test does not fix it — nextest isolates processes and they still share a filesystem path — so the slot is part of the address. ## Not in this commit The row's Done clause also asks that every existing `[[rule]]` row in this repository's own `batten.toml` carry a clean case. That is a much larger surface than the fixture corpus and is not here; what is here is the runner, the two cases the Ready block names, and the corpus cases that exercise both directions. Refs: CLOUD-313, CLOUD-310, CLOUD-283, CLOUD-418, CLOUD-63
…can read `Harness` names six hosts. The repository's own operating doctrine lived in `.claude/rules/`, which five of them cannot read — a completion gate shipping as a consumer-installable product, whose portability claim was true of its binary and false of its doctrine. A root `rules/` directory, with `.claude/rules/*.md` kept as pointer stubs. `rules/README.md` carries the decision, the two rejected alternatives and why, and the classification the row calls its real deliverable. The stubs are not a courtesy. They keep Claude Code's frontmatter `paths:` trigger — the one genuinely vendor-specific affordance in the whole surface, and a loading mechanism no neutral location has — and they keep the path alive for the ~17 governed `mise-tasks/*.sh` programs that cite it. Those are frozen under `shell-retirement`: editing one to follow the move is refused with one route and no override, so a move that broke their citations would have had no landable shape. A governed program's citation now resolves to a stub and is redirected in one hop. All five files are class (a) — mechanism-owned, the prose a pointer at a module, a task header or a source comment. **No rule is class (b).** That is the finding rather than an omission: nothing in these files described a Claude Code affordance, so the only vendor-specific thing in the surface was the loading mechanism. `scanning.md` carries the one class (c) residue and already declares it unowned in its own text — instrument suitability has no honest exit code, and rule 3 refuses a gate over a judgement. CLOUD-1152 forbade relocating prose ahead of the channel meant to carry it, and was `blockedBy` CLOUD-1362 for it. That is Done: `AdvisoryReach.delivered_on` is now non-empty for **2 of 6** hosts (`ClaudeCode`, `GeminiCli`), read from `hook.rs` rather than quoted. Relocation is not reach and the README says so; the remaining four are CLOUD-209's probe and CLOUD-44's shim. Three couplings, each found by a red test rather than by reading: - **`memories`' globs did not reach the new home.** A `mem:` edge written in a relocated rules file would have gone unjudged — the rule reports green over a file it never selects. `rules/*.md` and `rules/**/*.md` join its `line_sources`, and the fixture gains the same reach so the tier still discriminates. - **Two fixtures had their own paths rewritten under them.** `contract_drift` and `rules_drift` each declare a config naming `.claude/rules/**` and then write a file into it; a blanket citation re-path moved the write and left the config, which is fixture data rather than a citation. Reverted in both. - **An AGENTS.md trim evicted a load-bearing pointer.** Paying for the new routing row by shortening rule 8 dropped the CLOUD-591 citation, and `identity_precedence` caught it — CLOUD-994's class, live, in the same commit that cites it. The citations are restored and the line is paid for out of rule 7's restatement instead. `batten policy budget` was measured before and after: 199 lines and ~3446 tokens before, over at 200/3502 with the routing row added, clean again once the restatement came out. The budget was not widened. Every relocated rule is reachable from `AGENTS.md` in one hop — the routing table names all five, `policy-modules.md` included, which it did not before. Nothing was dropped in transit: the five files are `git mv`'d, so the diff shows renames, and the only content change is that each file's frontmatter went to its stub. `.serena/memories/**` is a second vendor surface with its own loading contract. §2 puts it out of scope and `rules/README.md` names it as inherited debt rather than folding it in. Refs: CLOUD-1152, CLOUD-1362, CLOUD-1150, CLOUD-994, CLOUD-1119, CLOUD-683 Admits: cf5a2a651a4150a96a451d9e5690538ec017a78e1ab71e531fabe66e018073e0 Admits-rule: protected-mutation Admits-verdict: path write refused Admits-subject: batten.toml Admits-anchor: call:1cc305dbd9fb59a13ea2b21d3407cfb8ce816a8a Admits-epoch: b8a9dba6792f8841cdee7b69cd1ea14c2392ba374dc05b6aff154b8b032e3115 Admits-author: alec@wenzowski.com Admits-prev: 1f335ff6f89a7f7f13207996572b4913ddef80e9e0f7cc9629b0ff46f80531e1 Admits-answer-lost: The relocation would be silent where it matters most. `memories` selects `.claude/*.md` and `.claude/**/*.md` and nothing else, so a `mem:` reference written in a relocated rules file reaches no glob and the rule reports green over a file it never read — the dead-gate shape this repository keeps re-meeting, caught here by a red fixture rather than by reading. `rules-drift` would likewise stop holding the key lists it exists to hold, and every `document` route would send a reader to a stub instead of to the rule. Admits-answer-precondition: `[[rule]]` rows and their `line_sources` are declared in batten.toml and nowhere else, so no other surface can express this. The edit follows the doctrine to its new home: `rules-drift` and `pr-partition-restated` re-point at `rules/*.md`, `memories` gains `rules/*.md` and `rules/**/*.md` so a `mem:` edge written there is still judged, `[contract] tracked` watches both the new home and the stubs, and every `[[verdict.route]]` document target names the authority rather than a pointer. A reviewer reads each as a path change in the diff it lands in. Admits-answer-rejected-route: config read first does not apply: the file has been read, and the finding IS that its globs and route targets name a directory the doctrine has left. patch run first does not apply either — there is no patch surface for a `line_sources` entry or a `[[verdict.route]]` target; the row IS the config.
…rry them CLOUD-1116 asks for a protocol before a measurement, and for a gate over the record set rather than over the treatment. Both halves land here. `bench/agentic/method.toml` owns what "improved" means: the attribution window, what is held constant, the seven measured outcomes, the four dimensions deliberately NOT measured, and the adjudication — paired difference, deterministic only, `blocking_on_model_judgement = false`. `bench/agentic/trials.toml` owns six paired trials, three on gate feedback and three on companion skills. **Grounded in this repository, not schematic.** Every `evidence` names a real `[[verdict]]` token, every `fixture` is a real commit on `origin/main`, and every companion skill is one this repository ships. A trial written against an invented refusal measures a prompt, not a gate. The registry was re-measured rather than quoted: **158 `[[verdict]]` rows and 195 `[[verdict.route]]` entries** — the figure CLOUD-1116 carries as a correction (52/66) is itself stale, as was the 37/49 that one corrected. **Every row is `pending`.** These are declared trials, not results. **A separate record set from `bench/tokens/`, and that file's own invariant is why.** Every arm there is run `runs` times and compared BYTE-FOR-BYTE; a model arm is never byte-stable, so a model row either breaks the invariant or is excluded from the comparison — leaving `token-bench-check`'s drift gate deciding nothing about exactly the rows this work adds. Nothing here edits it. `agentic-experiment-record`, read-only, `scope = "tree"`, three predicates: - **`agentic-record-incomplete`** — a trial naming fewer than the ten top-level keys, both arms and the falsifier's three fields; or a method record naming no window, no held-constant list, no measured/unmeasured split, or no adjudication. - **`agentic-finding-unsupported`** — a disposition asserting a finding with no `[trial.result]`, or one the method record never declared. `pending` asserts nothing and is the one value that owes no result. This is CLOUD-680's laundering shape arriving through the record channel: the cheapest route to a finding is to write one down. - **`agentic-record-unreadable`** — the could-not-look arm over both paths. It decides completeness and nothing else. Whether the `class` prose HELPED an agent is a judgement, and non-negotiable rule 3 keeps a judgement out of a verdict. Why a gate at all, since nobody is forced to write a trial: the failure mode is not an absent record, it is a HALF one. CLOUD-1089 paid for that twice in one survey — jscpd reporting "0 files analyzed", sonarjs with an undefined parser, both exit 0, both would have been written up as clean results. A trial missing its baseline arm reads exactly like a trial that ran. - **`sources` was the wrong column, and the could-not-look arm was dead.** `sources` is the GLOB field: an absent record matches nothing, is never declared, is never acquired, and never reaches `input.tree.missing`. Over the compiled binary a tree carrying NEITHER record answered nothing at all — a deleted `method.toml` and a satisfied one, byte-identical on the decision surface. `documents` is the literal field and is unioned in unconditionally. Every other case in the file passed over the broken spelling. - **`method.toml`'s `attribution_window` was unparseable TOML** — a basic string across two lines. The module reported it as `fixture-missing … (unparsed)` on its first run, which is the channel doing its job before a human looked. - **`object.union` merges recursively**, so two load-time cases meaning to REPLACE a sub-table were keeping the key they removed and passing for the wrong reason. `replacing` composes remove-then-union; the module says so at its site. A fourth, from the mediated boundary rather than a test: `git_in(&dir, ["init"])` is a hand-rolled fork `fixture-forks` refuses, and `common::init_repo` is the template copy it exists to protect. The row's acceptance clause requires the mutation replay over the record set to exist BEFORE the rule is set to `deny`, and it does — as a test rather than a paragraph, so it re-runs whenever the record set grows. Each required key is removed from each committed trial in turn and the gate must fire. Measured 2026-09-05: **6 trial rows, 13 required keys, 78 mutations, 78 fired, 0 false positives**, where the zero is `the_committed_records_satisfy_this_gate` asserting the unmutated set decides nothing. Two `#MUTANT` rows join `$MUTANT_GATES`, both anchored on the leading tab so the script reaches the rule body rather than its own declaration — a script matching its own row is one edit away from CLOUD-1445's inert mutation. ## The prune floor, refreshed because this bundle is what moved it `target-prune` refused `verify` on `[prune.warm]` measured against a tree that no longer exists: basis `crates/batten/tests/**/*.rs`, declared 197, live 208, tolerance 10. This bundle's own two suites are part of that 11, so refreshing it here rather than filing it is the honest place. Both floors scale by the stem model the block already states — warm 45.54/stem x 208 = 9472, cold 108.91 x 208 = 22653 — and neither is an independent measurement, which the ledger paragraph says in those words. Refs: CLOUD-1116, CLOUD-1089, CLOUD-680, CLOUD-1445 Admits: b5b2f6da78d6ac648923198e0d3f1cc0283fe4ce68f2d4c5276f934fa35e445a Admits-rule: protected-mutation Admits-verdict: path write refused Admits-subject: batten.toml Admits-anchor: call:7e63418e0fb8719875b1b825a5540d370abd50d5 Admits-epoch: b2133eee53e6e7ca7b4aef0128952865547c81b1e6d2d5f5b86302610711fc0c Admits-author: alec@wenzowski.com Admits-prev: - Admits-answer-lost: The branch cannot land at all: `target-prune` is on `verify`'s path and refuses on the stale basis, so every gate downstream of it is unreached. The gate's own reasoning names what the staleness costs if left: retained bytes are `keep x stems x size`, so a floor taken against a smaller stem count PASSES and then lets the build write more than it budgeted for, and the exhaustion arrives as a rustc IO error inside a test run rather than as a disk fault. This branch is 11 stems past the basis, one outside the tolerance, which is the gate working exactly as its own paragraph describes. Admits-answer-precondition: The floors, their `measured` dates and the two `[prune.*.basis]` counts are declared in batten.toml and nowhere else, so no other surface can express this. `target-prune` refused `verify` naming exactly this: `[prune.warm] was measured against a tree that no longer exists — basis crates/batten/tests/**/*.rs, declared 197, live 208, tolerance 10`, and its own remedy sentence is "Re-measure the floor and move `count` and `measured` together". The write moves `count` 197 -> 208 in both basis blocks, scales both floors by the stem model this block already states (warm 8971/197 = 45.54 per stem, x208 = 9472; cold 21455/197 = 108.91, x208 = 22653), sets `measured` to 2026-09-05, and adds the ledger paragraph recording the move and that neither floor is an independent measurement. A reviewer reads it as a floor refresh in the diff it lands in. Admits-answer-rejected-route: config read first does not apply: the block has been read, and the finding IS that its `count` and `measured` name a tree that no longer exists. patch run first does not apply either -- there is no patch surface for a `[prune.warm]` floor or a `[prune.warm.basis]` count; the rows ARE the config. Weakens: rule-predicate-changed rule[claim-order-is-stated].line_sources Weakens: rule-predicate-changed rule[hk-fix-selection].line_sources Weakens: rule-predicate-changed rule[memory-graph].line_sources Weakens: rule-predicate-changed rule[pr-partition-restated].line_sources Weakens: rule-predicate-changed rule[rules-drift].line_sources Weakens: rule-predicate-changed rule[no-tracker-key-in-core].regex
5b06ff5 to
fca9130
Compare
|
❌ The last analysis has failed. |
Closes CLOUD-761
Closes CLOUD-611
Closes CLOUD-275
Closes CLOUD-313
Closes CLOUD-1152
Closes CLOUD-1116
The
campaign-recordbundle: six rows, six commits, one branch. Each commit stands on its own and is reviewable on its own; this body is the map.CLOUD-761 — one conformance set for an issue key
CLOUD-1142 landed the definition —
ready::Grammar::key_ofandkeys_inown the three axes (case sensitive, the explicit character-class boundary, the project prefix mandatory). What it could not land is the obligation: nineteen governed shell consumers still re-derive the key, each retires whole under CLOUD-1164, and nothing made a successor's tier import the examples rather than re-type them. Re-typing is how a twentieth derivation arrives, and it arrives green.crates/batten/tests/it/issue_key.rsexportsconformance()— the row's fixed seven, each a site that behaves differently today:CLOUD-757cloud-757AB-1,Z-9,A-1fooCLOUD-1xCLOUD-179CLOUD-17is not inside itBoth questions are asked, because they are different ones and the four shell
caseglobs conflate them —[A-Z]*-[0-9]*accepts all three prefix near-misses because a glob cannot anchor. The reader under test is resolved from the committed[[pattern]]table, not a fixture expression: a fixture would let the registry row change while every case kept passing.no-tracker-key-in-corecarried one spelling,CLOUD-\[0-9\]— the way the twenty measured derivations happened to be written, and nothing else. A class a second spelling escapes is not a class. It is an alternation now, andno-tracker-key-in-modulesis its twin overpolicy/**, which had no gate at all. Measured over the tree: zero hits under either glob for the added spellings, so this adds a refusal and removes none.CLOUD-611 —
lock-tool-missingasks the presence question of every backendlock-completeasks two questions and one exemption was wired into both. "Does this entry lock a url and a checksum?" is genuinely inapplicable tonpm:,pipx:,cargo:,go:,gem:,core:*. "Does this tool have a lockfile entry at all?" applies to every backend, becausemise install --lockeddemands a row whatever the row contains.Measured on CLOUD-580's branch:
mise.tomlgained apipx:pin with nomise.lockrow;lock-completeexited 0,verifywas green twice on two SHAs, and every mise-action job died at the install step —zizmorandcommit-lintincluded, because--lockedvalidates the whole file rather than the job's install list. That is CLOUD-333's own motivating incident reproduced verbatim, in a backend the clause it added exempts, one issue after it closed.The conjunct is gone.
locks_nothingstays where it belongs — on the url question, read off a row that EXISTS.CLOUD-275 —
sha -> {model, harness, session}over the decision logdecision::attribution_forgroups the guard-decision log by (model, harness, session, dirty) and holds each record's ownCaller. The declared mutation is worth a sentence because the first one survived: it mutatedProvenance::from_host, andCaller::undeclared()constructsProvenance::Unknowndirectly and never calls it. Moved ontoas_str, where it discriminates.CLOUD-313 — score a rule's cases as reported / validated / noisy / missing
fixture_repos.rsgrowsOutcome,cases_in(),reported_pointers()andscore(), plusmaterialize_intoto fix a scratch-dir race the scoring exposed.CLOUD-1152 — the doctrine moves to a home five more harnesses can read
Harnessnames six hosts. The repository's own operating doctrine lived in.claude/rules/, which five of them cannot read — a completion gate shipping as a consumer-installable product, whose portability claim was true of its binary and false of its doctrine.Root
rules/, with.claude/rules/*.mdkept as pointer stubs. The stubs are not a courtesy: they keep Claude Code's frontmatterpaths:trigger — the one genuinely vendor-specific affordance in the surface — and they keep the path alive for the ~17 governedmise-tasks/*.shprograms that cite it, which are frozen undershell-retirement.rules/README.mdcarries the decision, the rejected alternatives, and the classification the row calls its real deliverable: all five files are class (a), and no rule is class (b) — nothing in them described a Claude Code affordance, so the only vendor-specific thing in the whole surface was the loading mechanism.Three couplings, each found by a red test rather than by reading:
memories' globs did not reach the new home (amem:edge written there would have gone unjudged); two fixtures had their own config paths rewritten under them; and an AGENTS.md trim evicted the CLOUD-591 citation — CLOUD-994's class, live, in the same commit that cites it.CLOUD-1116 — the agentic trials, and a gate over the records
bench/agentic/method.tomlowns what "improved" means — attribution window, what is held constant, seven measured outcomes, four dimensions deliberately not measured, and adjudication that is deterministic only (blocking_on_model_judgement = false).bench/agentic/trials.tomldeclares six paired trials, three on gate feedback and three on companion skills. Everyevidencenames a real[[verdict]]token, everyfixtureis a real commit onorigin/main, and every row ispending— these are declared trials, not results.A separate record set from
bench/tokens/, and that file's own invariant is why: every arm there is compared BYTE-FOR-BYTE across runs, which no model arm can satisfy, so a model row would either break the invariant or be excluded from the comparison — leavingtoken-bench-check's drift gate deciding nothing about exactly the rows this adds.agentic-experiment-recorddecides completeness over both records and nothing else. Whether a treatment HELPED is a judgement, and non-negotiable rule 3 keeps a judgement out of a verdict. The disposition clause is the one that will fire in anger:pendingasserts nothing, and every other disposition owes a[trial.result]— CLOUD-680's laundering shape arriving through the record channel, where the cheapest route to a finding is to write one down.severity = "deny"is licensed by the row's own acceptance clause: the mutation replay exists, as a test rather than a paragraph, so it re-runs whenever the record set grows. 6 trial rows, 13 required keys, 78 mutations, 78 fired, 0 false positives.Three things the second tier caught that reading did not:
sourceswas the glob column where the literal one was needed, so the could-not-look arm was dead and a tree missing both records answered nothing at all;method.toml'sattribution_windowwas unparseable TOML, reported by the engine asfixture-missing … (unparsed)on its first run; andobject.unionmerges recursively, so two load-time cases meant to REPLACE a sub-table were passing for the wrong reason. A fourth, fromopa check -srather than a test: a helper parameter namedobjectshadowed theobject.*builtin namespace — regorus resolved through it and ran 22 cases green, and the type checker refuses it, which is a gate whose behaviour depends on which binary reads it.Not in this PR
CLOUD-1174 is back in Backlog with a measured comment. Its §2 unit discriminator is stale — CLOUD-1149/1219/1224 inverted it —
repoints_at_the_declared_successorneeds a successor that does not exist at census time, and the home column cannot be gated because no live program declares a home. Building it as specified would have shipped a census that measures the wrong thing.Disclosure
Six
rule-predicate-changedsmells are groomed and admitted. Five re-pathline_sourcestorules/*.md, following the files (memory-graphadditionally GAINSrules/*.mdandrules/**/*.md); the sixth widensno-tracker-key-in-core's regex, which strictly adds refusals. The clauses were groomed onto CLOUD-1152 and CLOUD-761 on 2026-09-05 and not before the work, on an explicit owner override — both Ready blocks say so in those words, and the trailers on the last commit name the same six pairs.[prune.warm]/[prune.cold]are refreshed 197 → 208 stems because this bundle's own two suites are part of what moved them. Both floors scale by the stem model the block already states; neither is an independent measurement, which the ledger paragraph says explicitly.Refs: CLOUD-1142, CLOUD-1164, CLOUD-418, CLOUD-1369, CLOUD-333, CLOUD-580, CLOUD-1362, CLOUD-994, CLOUD-1089, CLOUD-680, CLOUD-1445, CLOUD-1174
Generated by Claude Code