Skip to content

feat(bundle): the campaign-record bundle — one conformance set, one presence question, and a record set that has to be checkable - #870

Draft
wenzowski wants to merge 6 commits into
mainfrom
claude/cloud-761-issue-key-conformance-fixture
Draft

feat(bundle): the campaign-record bundle — one conformance set, one presence question, and a record set that has to be checkable#870
wenzowski wants to merge 6 commits into
mainfrom
claude/cloud-761-issue-key-conformance-fixture

Conversation

@wenzowski

@wenzowski wenzowski commented Sep 5, 2026

Copy link
Copy Markdown
Contributor

Closes CLOUD-761
Closes CLOUD-611
Closes CLOUD-275
Closes CLOUD-313
Closes CLOUD-1152
Closes CLOUD-1116

The campaign-record bundle: six rows, six commits, one branch. Each commit stands on its own and is reviewable on its own; this body is the map.

CLOUD-761 — one conformance set for an issue key

CLOUD-1142 landed the definitionready::Grammar::key_of and keys_in own the three axes (case sensitive, the explicit character-class boundary, the project prefix mandatory). What it could not land is the obligation: nineteen governed shell consumers still re-derive the key, each retires whole under CLOUD-1164, and nothing made a successor's tier import the examples rather than re-type them. Re-typing is how a twentieth derivation arrives, and it arrives green.

crates/batten/tests/it/issue_key.rs exports conformance() — the row's fixed seven, each a site that behaves differently today:

subject whole key contains
CLOUD-757 yes 1
cloud-757 no 0
AB-1, Z-9, A-1foo no 0
CLOUD-1x no, and it carries one 1
CLOUD-179 yes, and CLOUD-17 is not inside it 1

Both questions are asked, because they are different ones and the four shell case globs conflate them — [A-Z]*-[0-9]* accepts all three prefix near-misses because a glob cannot anchor. The reader under test is resolved from the committed [[pattern]] table, not a fixture expression: a fixture would let the registry row change while every case kept passing.

no-tracker-key-in-core carried one spelling, CLOUD-\[0-9\] — the way the twenty measured derivations happened to be written, and nothing else. A class a second spelling escapes is not a class. It is an alternation now, and no-tracker-key-in-modules is its twin over policy/**, which had no gate at all. Measured over the tree: zero hits under either glob for the added spellings, so this adds a refusal and removes none.

CLOUD-611 — lock-tool-missing asks the presence question of every backend

lock-complete asks two questions and one exemption was wired into both. "Does this entry lock a url and a checksum?" is genuinely inapplicable to npm:, pipx:, cargo:, go:, gem:, core:*. "Does this tool have a lockfile entry at all?" applies to every backend, because mise install --locked demands a row whatever the row contains.

Measured on CLOUD-580's branch: mise.toml gained a pipx: pin with no mise.lock row; lock-complete exited 0, verify was green twice on two SHAs, and every mise-action job died at the install step — zizmor and commit-lint included, because --locked validates the whole file rather than the job's install list. That is CLOUD-333's own motivating incident reproduced verbatim, in a backend the clause it added exempts, one issue after it closed.

The conjunct is gone. locks_nothing stays where it belongs — on the url question, read off a row that EXISTS.

CLOUD-275 — sha -> {model, harness, session} over the decision log

decision::attribution_for groups the guard-decision log by (model, harness, session, dirty) and holds each record's own Caller. The declared mutation is worth a sentence because the first one survived: it mutated Provenance::from_host, and Caller::undeclared() constructs Provenance::Unknown directly and never calls it. Moved onto as_str, where it discriminates.

CLOUD-313 — score a rule's cases as reported / validated / noisy / missing

fixture_repos.rs grows Outcome, cases_in(), reported_pointers() and score(), plus materialize_into to fix a scratch-dir race the scoring exposed.

CLOUD-1152 — the doctrine moves to a home five more harnesses can read

Harness names six hosts. The repository's own operating doctrine lived in .claude/rules/, which five of them cannot read — a completion gate shipping as a consumer-installable product, whose portability claim was true of its binary and false of its doctrine.

Root rules/, with .claude/rules/*.md kept as pointer stubs. The stubs are not a courtesy: they keep Claude Code's frontmatter paths: trigger — the one genuinely vendor-specific affordance in the surface — and they keep the path alive for the ~17 governed mise-tasks/*.sh programs that cite it, which are frozen under shell-retirement. rules/README.md carries the decision, the rejected alternatives, and the classification the row calls its real deliverable: all five files are class (a), and no rule is class (b) — nothing in them described a Claude Code affordance, so the only vendor-specific thing in the whole surface was the loading mechanism.

Three couplings, each found by a red test rather than by reading: memories' globs did not reach the new home (a mem: edge written there would have gone unjudged); two fixtures had their own config paths rewritten under them; and an AGENTS.md trim evicted the CLOUD-591 citation — CLOUD-994's class, live, in the same commit that cites it.

CLOUD-1116 — the agentic trials, and a gate over the records

bench/agentic/method.toml owns what "improved" means — attribution window, what is held constant, seven measured outcomes, four dimensions deliberately not measured, and adjudication that is deterministic only (blocking_on_model_judgement = false). bench/agentic/trials.toml declares six paired trials, three on gate feedback and three on companion skills. Every evidence names a real [[verdict]] token, every fixture is a real commit on origin/main, and every row is pending — these are declared trials, not results.

A separate record set from bench/tokens/, and that file's own invariant is why: every arm there is compared BYTE-FOR-BYTE across runs, which no model arm can satisfy, so a model row would either break the invariant or be excluded from the comparison — leaving token-bench-check's drift gate deciding nothing about exactly the rows this adds.

agentic-experiment-record decides completeness over both records and nothing else. Whether a treatment HELPED is a judgement, and non-negotiable rule 3 keeps a judgement out of a verdict. The disposition clause is the one that will fire in anger: pending asserts nothing, and every other disposition owes a [trial.result] — CLOUD-680's laundering shape arriving through the record channel, where the cheapest route to a finding is to write one down.

severity = "deny" is licensed by the row's own acceptance clause: the mutation replay exists, as a test rather than a paragraph, so it re-runs whenever the record set grows. 6 trial rows, 13 required keys, 78 mutations, 78 fired, 0 false positives.

Three things the second tier caught that reading did not: sources was the glob column where the literal one was needed, so the could-not-look arm was dead and a tree missing both records answered nothing at all; method.toml's attribution_window was unparseable TOML, reported by the engine as fixture-missing … (unparsed) on its first run; and object.union merges recursively, so two load-time cases meant to REPLACE a sub-table were passing for the wrong reason. A fourth, from opa check -s rather than a test: a helper parameter named object shadowed the object.* builtin namespace — regorus resolved through it and ran 22 cases green, and the type checker refuses it, which is a gate whose behaviour depends on which binary reads it.

Not in this PR

CLOUD-1174 is back in Backlog with a measured comment. Its §2 unit discriminator is stale — CLOUD-1149/1219/1224 inverted it — repoints_at_the_declared_successor needs a successor that does not exist at census time, and the home column cannot be gated because no live program declares a home. Building it as specified would have shipped a census that measures the wrong thing.

Disclosure

Six rule-predicate-changed smells are groomed and admitted. Five re-path line_sources to rules/*.md, following the files (memory-graph additionally GAINS rules/*.md and rules/**/*.md); the sixth widens no-tracker-key-in-core's regex, which strictly adds refusals. The clauses were groomed onto CLOUD-1152 and CLOUD-761 on 2026-09-05 and not before the work, on an explicit owner override — both Ready blocks say so in those words, and the trailers on the last commit name the same six pairs.

[prune.warm]/[prune.cold] are refreshed 197 → 208 stems because this bundle's own two suites are part of what moved them. Both floors scale by the stem model the block already states; neither is an independent measurement, which the ledger paragraph says explicitly.

Refs: CLOUD-1142, CLOUD-1164, CLOUD-418, CLOUD-1369, CLOUD-333, CLOUD-580, CLOUD-1362, CLOUD-994, CLOUD-1089, CLOUD-680, CLOUD-1445, CLOUD-1174


Generated by Claude Code

@coderabbitai

coderabbitai Bot commented Sep 5, 2026

Copy link
Copy Markdown

Important

Review skipped

Too many files!

This PR contains 193 files, which is 93 over the limit of 100.

To get a review, reduce the PR to 100 files or fewer by splitting it into smaller PRs or changing its base branch.

Upgrade to a paid plan to raise the limit.

This review couldn't start because sufficient usage credits or metered capacity aren't available. Add credits or update usage-based reviews in the billing tab, then retry.

Check out review usage here.

⚙️ Run configuration

Configuration used: Path: .coderabbit.yaml

Review profile: CHILL

Plan: Team

Run ID: a2154ceb-5c95-436f-8bf0-03065a25a8b9

📥 Commits

Reviewing files that changed from the base of the PR and between 92ef45c and fca9130.

⛔ Files ignored due to path filters (1)
  • hk.pkl is excluded by !**/*.pkl
📒 Files selected for processing (193)
  • .claude/rules/commits.md
  • .claude/rules/policy-modules.md
  • .claude/rules/rust.md
  • .claude/rules/scanning.md
  • .claude/rules/toolchain.md
  • .config/nextest.toml
  • .serena/project.yml
  • AGENTS.md
  • Cargo.toml
  • batten.toml
  • bench/agentic/method.toml
  • bench/agentic/trials.toml
  • bench/gates/RESULTS.md
  • bench/rust-suites/RESULTS.md
  • bench/tokens/RESULTS.md
  • bench/tokens/method.toml
  • crates/batten/src/admission.rs
  • crates/batten/src/capture.rs
  • crates/batten/src/captured.rs
  • crates/batten/src/claim.rs
  • crates/batten/src/decision.rs
  • crates/batten/src/doctor.rs
  • crates/batten/src/environment.rs
  • crates/batten/src/exec.rs
  • crates/batten/src/facts.rs
  • crates/batten/src/fetch.rs
  • crates/batten/src/forge.rs
  • crates/batten/src/git.rs
  • crates/batten/src/gitwrite.rs
  • crates/batten/src/handler.rs
  • crates/batten/src/hook.rs
  • crates/batten/src/land.rs
  • crates/batten/src/landed.rs
  • crates/batten/src/lease.rs
  • crates/batten/src/lib.rs
  • crates/batten/src/mcp.rs
  • crates/batten/src/mutate.rs
  • crates/batten/src/pinned.rs
  • crates/batten/src/policy.rs
  • crates/batten/src/preset.rs
  • crates/batten/src/ready.rs
  • crates/batten/src/receipt.rs
  • crates/batten/src/recorder.rs
  • crates/batten/src/refusal.rs
  • crates/batten/src/review.rs
  • crates/batten/src/rules.rs
  • crates/batten/src/secrets.rs
  • crates/batten/src/surface.rs
  • crates/batten/src/symbols.rs
  • crates/batten/src/task.rs
  • crates/batten/src/transcript.rs
  • crates/batten/src/verdict.rs
  • crates/batten/src/wiring.rs
  • crates/batten/tests/fixtures/briefs/complete.md
  • crates/batten/tests/fixtures/briefs/missing-check.md
  • crates/batten/tests/fixtures/briefs/unrunnable-check.md
  • crates/batten/tests/fixtures/repos/forbid-deny/expected.in
  • crates/batten/tests/fixtures/repos/forbid-deny/lib.rs.in
  • crates/batten/tests/fixtures/repos/forbid-exclude-comment/run.sh.in
  • crates/batten/tests/fixtures/repos/forbid-quote-load-bearing/batten.toml.in
  • crates/batten/tests/fixtures/repos/forbid-quote-load-bearing/expected.in
  • crates/batten/tests/fixtures/repos/forbid-quote-load-bearing/pins.toml.in
  • crates/batten/tests/fixtures/transcripts/memory-injection-unsupported.jsonl.in
  • crates/batten/tests/fixtures/transcripts/memory-injection.jsonl.in
  • crates/batten/tests/it/acquisition_metric.rs
  • crates/batten/tests/it/acquisition_sweep.rs
  • crates/batten/tests/it/agentic_record.rs
  • crates/batten/tests/it/ambient_authority.rs
  • crates/batten/tests/it/attribution_provenance.rs
  • crates/batten/tests/it/bats_invocation.rs
  • crates/batten/tests/it/board_record.rs
  • crates/batten/tests/it/board_state_claim.rs
  • crates/batten/tests/it/bot_lane.rs
  • crates/batten/tests/it/call_arguments.rs
  • crates/batten/tests/it/call_ceiling.rs
  • crates/batten/tests/it/capture_fidelity.rs
  • crates/batten/tests/it/ci_cache_declared.rs
  • crates/batten/tests/it/ci_hygiene.rs
  • crates/batten/tests/it/ci_parity.rs
  • crates/batten/tests/it/ci_suite_lane.rs
  • crates/batten/tests/it/claim_carry.rs
  • crates/batten/tests/it/claim_order.rs
  • crates/batten/tests/it/cli.rs
  • crates/batten/tests/it/commit_admission.rs
  • crates/batten/tests/it/commit_arm_sequencing.rs
  • crates/batten/tests/it/common/mod.rs
  • crates/batten/tests/it/config_authority_boundary.rs
  • crates/batten/tests/it/connector_allow_door.rs
  • crates/batten/tests/it/connector_not_granted.rs
  • crates/batten/tests/it/container_health.rs
  • crates/batten/tests/it/contract_drift.rs
  • crates/batten/tests/it/doctor.rs
  • crates/batten/tests/it/document_read_count.rs
  • crates/batten/tests/it/egress_fencing.rs
  • crates/batten/tests/it/filed_here.rs
  • crates/batten/tests/it/fixture_forks.rs
  • crates/batten/tests/it/fixture_repos.rs
  • crates/batten/tests/it/fuzz_corpus.rs
  • crates/batten/tests/it/gh_guard.rs
  • crates/batten/tests/it/git_facts.rs
  • crates/batten/tests/it/harness_grant.rs
  • crates/batten/tests/it/hk_fix_selection.rs
  • crates/batten/tests/it/hook_cost.rs
  • crates/batten/tests/it/hook_profile.rs
  • crates/batten/tests/it/hook_skip_local.rs
  • crates/batten/tests/it/identity_precedence.rs
  • crates/batten/tests/it/issue_key.rs
  • crates/batten/tests/it/landed_check.rs
  • crates/batten/tests/it/lock_complete.rs
  • crates/batten/tests/it/main.rs
  • crates/batten/tests/it/mcp_dispatch.rs
  • crates/batten/tests/it/mediated_admission.rs
  • crates/batten/tests/it/memories.rs
  • crates/batten/tests/it/memory_injection.rs
  • crates/batten/tests/it/mise_pin_agreement.rs
  • crates/batten/tests/it/narrow_adoption.rs
  • crates/batten/tests/it/perf_assert.rs
  • crates/batten/tests/it/perf_compare.rs
  • crates/batten/tests/it/pinned_programs.rs
  • crates/batten/tests/it/plan_complete.rs
  • crates/batten/tests/it/policy_input_narrowing.rs
  • crates/batten/tests/it/policy_presets.rs
  • crates/batten/tests/it/policy_tree.rs
  • crates/batten/tests/it/pr_partition_restated.rs
  • crates/batten/tests/it/preset_manifest.rs
  • crates/batten/tests/it/preset_segments.rs
  • crates/batten/tests/it/privileged_lane.rs
  • crates/batten/tests/it/process_group.rs
  • crates/batten/tests/it/ready.rs
  • crates/batten/tests/it/reclaim_report_once.rs
  • crates/batten/tests/it/release_provision_parity.rs
  • crates/batten/tests/it/retirement_doctrine.rs
  • crates/batten/tests/it/review_answered.rs
  • crates/batten/tests/it/rule_cost_census.rs
  • crates/batten/tests/it/rules_builtin_claims.rs
  • crates/batten/tests/it/rules_drift.rs
  • crates/batten/tests/it/run_shape_guard_door.rs
  • crates/batten/tests/it/sbom_inventory.rs
  • crates/batten/tests/it/scanner_taxonomy.rs
  • crates/batten/tests/it/session_provisioning.rs
  • crates/batten/tests/it/shell_retirement.rs
  • crates/batten/tests/it/shell_retirement_cost.rs
  • crates/batten/tests/it/shell_write_advisory.rs
  • crates/batten/tests/it/sinks.rs
  • crates/batten/tests/it/spawn_census.rs
  • crates/batten/tests/it/startup.rs
  • crates/batten/tests/it/stop_posture.rs
  • crates/batten/tests/it/submodule.rs
  • crates/batten/tests/it/symbols.rs
  • crates/batten/tests/it/target_prune.rs
  • crates/batten/tests/it/task_prose.rs
  • crates/batten/tests/it/task_registry.rs
  • crates/batten/tests/it/test_targets.rs
  • crates/batten/tests/it/todo_promotion.rs
  • crates/batten/tests/it/tool_verdict_facts.rs
  • crates/batten/tests/it/verdict_registry.rs
  • crates/batten/tests/it/wiring_reclaim.rs
  • crates/batten/tests/it/worktree_registration.rs
  • crates/batten/tests/policy_modules.rs
  • mise.toml
  • policy/agentic-experiment-record.rego
  • policy/claim-before-code.rego
  • policy/claim-order-is-stated.rego
  • policy/connector-not-granted.rego
  • policy/denials-outlive-the-turn.rego
  • policy/egress-fencing.rego
  • policy/fixture-forks.rego
  • policy/harness-grant.rego
  • policy/harness-wiring.rego
  • policy/hk-fix-selection.rego
  • policy/lock-complete.rego
  • policy/memories.rego
  • policy/obligations-bound.rego
  • policy/perf-assert.rego
  • policy/pr-partition-restated.rego
  • policy/release-provision-parity.rego
  • policy/rules-drift.rego
  • policy/run-shape.rego
  • policy/shell-retirement.rego
  • policy/shell-write-advisory.rego
  • policy/spawn-adapters.rego
  • policy/suite-subject-retirable.rego
  • policy/task-substitution.rego
  • policy/test-targets.rego
  • rules/README.md
  • rules/commits.md
  • rules/policy-modules.md
  • rules/rust.md
  • rules/scanning.md
  • rules/toolchain.md
  • schema/batten.schema.json
  • schema/policy-input.schema.json
  • skills/serena/SKILL.md

You can disable this status message by setting the reviews.review_status to false in the CodeRabbit configuration file.


Thanks for using CodeRabbit! It's free for OSS, and your support helps us grow. If you like it, consider giving us a shout-out.

❤️ Share

Comment @coderabbitai help to get the list of available commands.

@wenzowski
wenzowski force-pushed the claude/cloud-761-issue-key-conformance-fixture branch 2 times, most recently from 66fd221 to 5b06ff5 Compare September 5, 2026 19:48
@wenzowski wenzowski changed the title test(facts): hold every consumer to one conformance set for an issue key feat(bundle): the campaign-record bundle — one conformance set, one presence question, and a record set that has to be checkable Sep 5, 2026
CLOUD-1142 landed the definition — `ready::Grammar::key_of` and `keys_in`
own the three axes. What it could not land is the obligation: the nineteen
governed shell consumers still re-derive the key, each retires whole under
CLOUD-1164, and nothing made its successor's tier import the examples
rather than re-type them.

`crates/batten/tests/it/issue_key.rs` exports `conformance()` — the row's
fixed seven, decided by the engine's own reader resolved from the COMMITTED
`[[pattern]]` table, so a registry row's drift reddens the integration tier
and not only the crate's own unit tests.

`no-tracker-key-in-core` carried ONE spelling of the key expression, which
refuses the way the twenty measured derivations happened to be written and
waves through every other way of writing the same thing. It is an
alternation now, and `no-tracker-key-in-modules` is its twin over `policy/**`
— a `forbid` rather than a ratchet there, because that tree carries zero
derivations to wait out. Both are asserted by the committed-ruleset census,
whose full-stdout equality is what tells a mis-globbed row from a working one.

The declared mutation is bound in the test file, where `obligations-bound`
reads it, and declared again beside the reader in `ready.rs` — `engine-ready`
joins `$MUTANT_GATES`, so `mise run mutant` APPLIES it rather than only
binding it, closing the gap `commit_arm_sequencing.rs` records for its own row.

Refs: CLOUD-761, CLOUD-1142, CLOUD-1164, CLOUD-418

Admits: 1f335ff6f89a7f7f13207996572b4913ddef80e9e0f7cc9629b0ff46f80531e1
Admits-rule: protected-mutation
Admits-verdict: path write refused
Admits-subject: batten.toml
Admits-anchor: call:a16725c3350d90c01673dd1a521cb1e1f3e4057d
Admits-epoch: 55191757a873a6121f5968407711eb9d18456d12f8af05fd7850814f3823f9f6
Admits-author: alec@wenzowski.com
Admits-prev: b5b2f6da78d6ac648923198e0d3f1cc0283fe4ce68f2d4c5276f934fa35e445a
Admits-answer-lost: CLOUD-761's deliverable. `no-tracker-key-in-core` refuses exactly one spelling, the literal `CLOUD-[0-9]`, and `policy/**` carries no such gate at all — so a module or crate file composing `CLOUD-\d` is a second authority over the one definition of an issue key and nothing refuses it. A class a second spelling escapes is not a class; twenty derivations accumulated because nothing refused the twenty-first before it was typed.
Admits-answer-precondition: A `[[rule]]` row is declared in batten.toml and nowhere else, so no other surface can express this change. It adds one `forbid` row over `policy/**` (`no-tracker-key-in-modules`) and widens `no-tracker-key-in-core`'s regex from one spelling of the key expression to three; a reviewer reads both directly in the diff as one new row and one changed line.
Admits-answer-rejected-route: config read first does not apply: the file has been read, and the finding IS that the committed row's regex covers one spelling of three. patch run first does not apply either — there is no patch surface for a `[[rule]]` row, the row IS the config.
… locked ones

`lock-complete` asks two questions and one exemption was wired into both.
Whether an entry LOCKS a url is genuinely inapplicable to a backend that
resolves through its own package manager — that is what `locks_nothing` is
for. Whether a tool HAS an entry applies to every backend without exception,
because `mise install --locked` demands a row whatever the row contains.

`lock-tool-missing` inherited the exemption anyway, and its own comment called
that "a LENIENCE here rather than a necessity". Measured on CLOUD-580's branch,
in a backend the clause exempted and one issue after CLOUD-333 closed:
`"pipx:ntia-conformance-checker" = "5.0.3"` with no row passed at exit 0 and
`verify` green twice on two SHAs, then took every mise-action job red at the
install step — `zizmor` and `commit-lint` included, since `--locked` validates
the whole file rather than the job's install list.

The edit is one conjunct. `lock-tool-unlocked`'s url arm keeps the identical
call, and it is a genuinely different call: it reads the backend off a row that
EXISTS, where this clause has no row to read.

`key_backend` goes with it rather than being left dead — it existed only to
feed the exemption.

## The fixture that had been hiding behind the exemption

One load-time test needed changing, which the row named as the signal that a
removal reached the wrong clause. It did not, and the reason is worth the
sentence: `test_a_backend_that_cannot_lock_is_exempt_from_locking_nothing`
keyed its lock entry `core:rust` while its manifest said `rust`, a pairing mise
does not produce — the committed `mise.lock` says `[[tools.rust]]` with
`backend = "core:rust"`. The presence lookup is by the lock's own key, so the
mismatch was always there and the exemption short-circuited before anything
could reach it. The fixture now spells the entry the way the tool writes it and
tests the url question alone.

Verified against the real tree as well as the fixtures: `batten check --rule
lock-complete` is silent over this repository, so the narrowing changed no
verdict on a correct tree — every exempt-backend pin here already carries a row.

The declared mutation re-creates the removed blind spot by re-exempting one
prefix on the presence clause; applied, it reddens only the new case and leaves
the seventeen others green.

Refs: CLOUD-611, CLOUD-333, CLOUD-580, CLOUD-418
…cision log

CLOUD-268 moved model, harness and session identifiers out of durable public
commit text, and `batten attribution` enforces that suppression over produced
commits. Suppression alone trades an unreliable public signal for nothing
unless the operational question stays answerable, so this is the read side:
`decision::attribution_for(repo_root, sha)`.

The join key and the payload were both already there — `Anchor.commit`, whose
own doc names this row, and `Caller`'s three fields, reserved by CLOUD-133.
What was missing is the query.

## A grouping, never a merge

A commit can carry many records, and they can disagree: two sessions, or a tree
that went dirty between two calls. Folding them into one row would invent a
fact no record carries, so each distinct `(caller, dirty)` pair is its own row
with the count that carried it. `dirty` is part of the key rather than a column
because it changes what the row claims — a record taken against a dirty tree
describes a call the commit does not fully contain.

## Absent is not unknown

No rows means the log holds nothing anchored to this sha: the could-not-look
case, and the ordinary state of a commit made outside a mediated session. A row
valued `unknown` means a record exists and the host exposed no identity.
Collapsing them makes a gap read as a fact — CLOUD-251's trap in this surface —
so both halves are asserted in one case, where neither can drift from the other.

The rows carry each record's OWN `Caller` rather than one rebuilt from its
tokens. `Provenance::from_host` degrades empty and whitespace; it is
`Provenance`'s `Deserialize` that maps the `unknown` token back to `Unknown`. A
reconstruction through `from_host` would turn a degraded field into
`Declared("unknown")` — the query and the record disagreeing about what a
degraded field is, which is the one distinction this row exists to keep. Caught
while writing it, not by a test.

## Rule 4 structurally

`subject` and `context` are not fields of the answer, so no rendering choice can
leak them. The pointer-only case plants a real path and real context bytes in
the record and asserts none of it reaches the rendered answer, so it would fail
if a later author added such a field "for context".

## The declared mutation, and the one that survived first

`degraded-field-not-unknown` stops rendering the degraded token, so an
undeclared field serializes empty and reads back as `Declared("")`. Applied, it
reddens the degraded case alone and leaves the other three green.

Its first spelling mutated `Provenance::from_host` and SURVIVED:
`Caller::undeclared()` builds `Provenance::Unknown` directly and never calls
`from_host`, so the case never walked the mutated line. The file header records
that, because choosing a mutation over a path the case does not take is exactly
how a declaration becomes coverage that cannot fail.

Refs: CLOUD-275, CLOUD-133, CLOUD-268, CLOUD-251, CLOUD-418
…issing

A `[[rule]]` row is a classifier with two failure modes and the corpus observed
one. `expected.in` pins the whole of stdout, so a rule that stops firing does
change those bytes — but it fails as a stdout diff, which names the fixture and
not the rule, and it says nothing at all about a rule firing on a line nobody
meant it to.

The four instances in `batten.toml`'s own history are all that untested half: a
comment narrating a banned command, literals dropped because they fired on
ordinary prose, a leading quote carried purely to suppress a false positive, and
a self-match hazard nothing pinned. Measured (CLOUD-310): of 40 lines a literal
row reported across `mise-tasks/`, 8 were comments — 20% noise the suite could
not see.

## The expectation lives in the case

`# batten-case: <rule-id> clean|violating` declares the line that FOLLOWS it —
`#MUTANT`'s shape, for its reason: one authority, adjacent to its subject, with
no second file to keep in agreement. The three candidates the row listed each
fail on a rule this repository already holds, and the Ready block chose this
fourth.

Four outcomes, two of them named failures: `reported`, `validated`, `noisy`,
`missing`. The runner is `fixture_repos.rs` extended rather than a new program —
a new `mise-tasks/` file is born governed, so its own first bug fix would have
had nowhere to land. Output is the case's `path:line rule-id` and the outcome
token, never the matched bytes: a false-positive report whose subject is a
matched literal is exactly where a checker leaks what it was scanning.

## Shown able to fail, in the direction the row is about

`forbid-quote-load-bearing` reproduces `no-source-built-tool`'s shape without
depending on this repository's own rule table. Dropping the leading quote from
its pattern makes the rule report the comment explaining itself, and the suite
now fails as `pins.toml:3 no-built-backend noisy` rather than as a stdout diff.
The quote is pinned as load-bearing instead of remembered in a comment.

`forbid-exclude-comment`'s two comment lines were the historical false positive
and were clean only by someone's memory; they are declared cases now. The scorer
is negatively self-tested in both directions, and a finding one line off scores
as `missing` rather than as a pass.

## The self-match convention gains a test

Rule ids deliberately do not name their own literals, because a finding's
pointer output lands in a fixture inside the rule's own glob. That was a
convention with nothing behind it, and a marker naming a rule id is exactly what
a reintroduced self-match would break — so it is asserted over the corpus now.

## One defect this change caused and fixed

The scratch path `fixture-repos/<name>` was unambiguous with one consumer and I
added two more. Measured: the scoring test passed alone and scored `missing`
beside its sibling, reading a tree the other was still writing. Process-per-test
does not fix it — nextest isolates processes and they still share a filesystem
path — so the slot is part of the address.

## Not in this commit

The row's Done clause also asks that every existing `[[rule]]` row in this
repository's own `batten.toml` carry a clean case. That is a much larger surface
than the fixture corpus and is not here; what is here is the runner, the two
cases the Ready block names, and the corpus cases that exercise both directions.

Refs: CLOUD-313, CLOUD-310, CLOUD-283, CLOUD-418, CLOUD-63
…can read

`Harness` names six hosts. The repository's own operating doctrine lived in
`.claude/rules/`, which five of them cannot read — a completion gate shipping as
a consumer-installable product, whose portability claim was true of its binary
and false of its doctrine.

A root `rules/` directory, with `.claude/rules/*.md` kept as pointer stubs.
`rules/README.md` carries the decision, the two rejected alternatives and why,
and the classification the row calls its real deliverable.

The stubs are not a courtesy. They keep Claude Code's frontmatter `paths:`
trigger — the one genuinely vendor-specific affordance in the whole surface, and
a loading mechanism no neutral location has — and they keep the path alive for
the ~17 governed `mise-tasks/*.sh` programs that cite it. Those are frozen under
`shell-retirement`: editing one to follow the move is refused with one route and
no override, so a move that broke their citations would have had no landable
shape. A governed program's citation now resolves to a stub and is redirected in
one hop.

All five files are class (a) — mechanism-owned, the prose a pointer at a module,
a task header or a source comment. **No rule is class (b).** That is the finding
rather than an omission: nothing in these files described a Claude Code
affordance, so the only vendor-specific thing in the surface was the loading
mechanism. `scanning.md` carries the one class (c) residue and already declares
it unowned in its own text — instrument suitability has no honest exit code, and
rule 3 refuses a gate over a judgement.

CLOUD-1152 forbade relocating prose ahead of the channel meant to carry it, and
was `blockedBy` CLOUD-1362 for it. That is Done: `AdvisoryReach.delivered_on` is
now non-empty for **2 of 6** hosts (`ClaudeCode`, `GeminiCli`), read from
`hook.rs` rather than quoted. Relocation is not reach and the README says so; the
remaining four are CLOUD-209's probe and CLOUD-44's shim.

Three couplings, each found by a red test rather than by reading:

- **`memories`' globs did not reach the new home.** A `mem:` edge written in a
  relocated rules file would have gone unjudged — the rule reports green over a
  file it never selects. `rules/*.md` and `rules/**/*.md` join its `line_sources`,
  and the fixture gains the same reach so the tier still discriminates.
- **Two fixtures had their own paths rewritten under them.** `contract_drift` and
  `rules_drift` each declare a config naming `.claude/rules/**` and then write a
  file into it; a blanket citation re-path moved the write and left the config,
  which is fixture data rather than a citation. Reverted in both.
- **An AGENTS.md trim evicted a load-bearing pointer.** Paying for the new routing
  row by shortening rule 8 dropped the CLOUD-591 citation, and
  `identity_precedence` caught it — CLOUD-994's class, live, in the same commit
  that cites it. The citations are restored and the line is paid for out of rule
  7's restatement instead.

`batten policy budget` was measured before and after: 199 lines and ~3446 tokens
before, over at 200/3502 with the routing row added, clean again once the
restatement came out. The budget was not widened.

Every relocated rule is reachable from `AGENTS.md` in one hop — the routing table
names all five, `policy-modules.md` included, which it did not before. Nothing
was dropped in transit: the five files are `git mv`'d, so the diff shows renames,
and the only content change is that each file's frontmatter went to its stub.

`.serena/memories/**` is a second vendor surface with its own loading contract.
§2 puts it out of scope and `rules/README.md` names it as inherited debt rather
than folding it in.

Refs: CLOUD-1152, CLOUD-1362, CLOUD-1150, CLOUD-994, CLOUD-1119, CLOUD-683

Admits: cf5a2a651a4150a96a451d9e5690538ec017a78e1ab71e531fabe66e018073e0
Admits-rule: protected-mutation
Admits-verdict: path write refused
Admits-subject: batten.toml
Admits-anchor: call:1cc305dbd9fb59a13ea2b21d3407cfb8ce816a8a
Admits-epoch: b8a9dba6792f8841cdee7b69cd1ea14c2392ba374dc05b6aff154b8b032e3115
Admits-author: alec@wenzowski.com
Admits-prev: 1f335ff6f89a7f7f13207996572b4913ddef80e9e0f7cc9629b0ff46f80531e1
Admits-answer-lost: The relocation would be silent where it matters most. `memories` selects `.claude/*.md` and `.claude/**/*.md` and nothing else, so a `mem:` reference written in a relocated rules file reaches no glob and the rule reports green over a file it never read — the dead-gate shape this repository keeps re-meeting, caught here by a red fixture rather than by reading. `rules-drift` would likewise stop holding the key lists it exists to hold, and every `document` route would send a reader to a stub instead of to the rule.
Admits-answer-precondition: `[[rule]]` rows and their `line_sources` are declared in batten.toml and nowhere else, so no other surface can express this. The edit follows the doctrine to its new home: `rules-drift` and `pr-partition-restated` re-point at `rules/*.md`, `memories` gains `rules/*.md` and `rules/**/*.md` so a `mem:` edge written there is still judged, `[contract] tracked` watches both the new home and the stubs, and every `[[verdict.route]]` document target names the authority rather than a pointer. A reviewer reads each as a path change in the diff it lands in.
Admits-answer-rejected-route: config read first does not apply: the file has been read, and the finding IS that its globs and route targets name a directory the doctrine has left. patch run first does not apply either — there is no patch surface for a `line_sources` entry or a `[[verdict.route]]` target; the row IS the config.
…rry them

CLOUD-1116 asks for a protocol before a measurement, and for a gate over the
record set rather than over the treatment. Both halves land here.

`bench/agentic/method.toml` owns what "improved" means: the attribution window,
what is held constant, the seven measured outcomes, the four dimensions
deliberately NOT measured, and the adjudication — paired difference, deterministic
only, `blocking_on_model_judgement = false`. `bench/agentic/trials.toml` owns six
paired trials, three on gate feedback and three on companion skills.

**Grounded in this repository, not schematic.** Every `evidence` names a real
`[[verdict]]` token, every `fixture` is a real commit on `origin/main`, and every
companion skill is one this repository ships. A trial written against an invented
refusal measures a prompt, not a gate. The registry was re-measured rather than
quoted: **158 `[[verdict]]` rows and 195 `[[verdict.route]]` entries** — the
figure CLOUD-1116 carries as a correction (52/66) is itself stale, as was the
37/49 that one corrected.

**Every row is `pending`.** These are declared trials, not results.

**A separate record set from `bench/tokens/`, and that file's own invariant is
why.** Every arm there is run `runs` times and compared BYTE-FOR-BYTE; a model arm
is never byte-stable, so a model row either breaks the invariant or is excluded
from the comparison — leaving `token-bench-check`'s drift gate deciding nothing
about exactly the rows this work adds. Nothing here edits it.

`agentic-experiment-record`, read-only, `scope = "tree"`, three predicates:

- **`agentic-record-incomplete`** — a trial naming fewer than the ten top-level
  keys, both arms and the falsifier's three fields; or a method record naming no
  window, no held-constant list, no measured/unmeasured split, or no adjudication.
- **`agentic-finding-unsupported`** — a disposition asserting a finding with no
  `[trial.result]`, or one the method record never declared. `pending` asserts
  nothing and is the one value that owes no result. This is CLOUD-680's laundering
  shape arriving through the record channel: the cheapest route to a finding is to
  write one down.
- **`agentic-record-unreadable`** — the could-not-look arm over both paths.

It decides completeness and nothing else. Whether the `class` prose HELPED an
agent is a judgement, and non-negotiable rule 3 keeps a judgement out of a verdict.

Why a gate at all, since nobody is forced to write a trial: the failure mode is
not an absent record, it is a HALF one. CLOUD-1089 paid for that twice in one
survey — jscpd reporting "0 files analyzed", sonarjs with an undefined parser,
both exit 0, both would have been written up as clean results. A trial missing its
baseline arm reads exactly like a trial that ran.

- **`sources` was the wrong column, and the could-not-look arm was dead.**
  `sources` is the GLOB field: an absent record matches nothing, is never
  declared, is never acquired, and never reaches `input.tree.missing`. Over the
  compiled binary a tree carrying NEITHER record answered nothing at all — a
  deleted `method.toml` and a satisfied one, byte-identical on the decision
  surface. `documents` is the literal field and is unioned in unconditionally.
  Every other case in the file passed over the broken spelling.
- **`method.toml`'s `attribution_window` was unparseable TOML** — a basic string
  across two lines. The module reported it as `fixture-missing … (unparsed)` on
  its first run, which is the channel doing its job before a human looked.
- **`object.union` merges recursively**, so two load-time cases meaning to REPLACE
  a sub-table were keeping the key they removed and passing for the wrong reason.
  `replacing` composes remove-then-union; the module says so at its site.

A fourth, from the mediated boundary rather than a test: `git_in(&dir, ["init"])`
is a hand-rolled fork `fixture-forks` refuses, and `common::init_repo` is the
template copy it exists to protect.

The row's acceptance clause requires the mutation replay over the record set to
exist BEFORE the rule is set to `deny`, and it does — as a test rather than a
paragraph, so it re-runs whenever the record set grows. Each required key is
removed from each committed trial in turn and the gate must fire. Measured
2026-09-05: **6 trial rows, 13 required keys, 78 mutations, 78 fired, 0 false
positives**, where the zero is `the_committed_records_satisfy_this_gate` asserting
the unmutated set decides nothing.

Two `#MUTANT` rows join `$MUTANT_GATES`, both anchored on the leading tab so the
script reaches the rule body rather than its own declaration — a script matching
its own row is one edit away from CLOUD-1445's inert mutation.

## The prune floor, refreshed because this bundle is what moved it

`target-prune` refused `verify` on `[prune.warm]` measured against a tree that no
longer exists: basis `crates/batten/tests/**/*.rs`, declared 197, live 208,
tolerance 10. This bundle's own two suites are part of that 11, so refreshing it
here rather than filing it is the honest place. Both floors scale by the stem
model the block already states — warm 45.54/stem x 208 = 9472, cold 108.91 x 208
= 22653 — and neither is an independent measurement, which the ledger paragraph
says in those words.

Refs: CLOUD-1116, CLOUD-1089, CLOUD-680, CLOUD-1445

Admits: b5b2f6da78d6ac648923198e0d3f1cc0283fe4ce68f2d4c5276f934fa35e445a
Admits-rule: protected-mutation
Admits-verdict: path write refused
Admits-subject: batten.toml
Admits-anchor: call:7e63418e0fb8719875b1b825a5540d370abd50d5
Admits-epoch: b2133eee53e6e7ca7b4aef0128952865547c81b1e6d2d5f5b86302610711fc0c
Admits-author: alec@wenzowski.com
Admits-prev: -
Admits-answer-lost: The branch cannot land at all: `target-prune` is on `verify`'s path and refuses on the stale basis, so every gate downstream of it is unreached. The gate's own reasoning names what the staleness costs if left: retained bytes are `keep x stems x size`, so a floor taken against a smaller stem count PASSES and then lets the build write more than it budgeted for, and the exhaustion arrives as a rustc IO error inside a test run rather than as a disk fault. This branch is 11 stems past the basis, one outside the tolerance, which is the gate working exactly as its own paragraph describes.
Admits-answer-precondition: The floors, their `measured` dates and the two `[prune.*.basis]` counts are declared in batten.toml and nowhere else, so no other surface can express this. `target-prune` refused `verify` naming exactly this: `[prune.warm] was measured against a tree that no longer exists — basis crates/batten/tests/**/*.rs, declared 197, live 208, tolerance 10`, and its own remedy sentence is "Re-measure the floor and move `count` and `measured` together". The write moves `count` 197 -> 208 in both basis blocks, scales both floors by the stem model this block already states (warm 8971/197 = 45.54 per stem, x208 = 9472; cold 21455/197 = 108.91, x208 = 22653), sets `measured` to 2026-09-05, and adds the ledger paragraph recording the move and that neither floor is an independent measurement. A reviewer reads it as a floor refresh in the diff it lands in.
Admits-answer-rejected-route: config read first does not apply: the block has been read, and the finding IS that its `count` and `measured` name a tree that no longer exists. patch run first does not apply either -- there is no patch surface for a `[prune.warm]` floor or a `[prune.warm.basis]` count; the rows ARE the config.
Weakens: rule-predicate-changed rule[claim-order-is-stated].line_sources
Weakens: rule-predicate-changed rule[hk-fix-selection].line_sources
Weakens: rule-predicate-changed rule[memory-graph].line_sources
Weakens: rule-predicate-changed rule[pr-partition-restated].line_sources
Weakens: rule-predicate-changed rule[rules-drift].line_sources
Weakens: rule-predicate-changed rule[no-tracker-key-in-core].regex
@wenzowski
wenzowski force-pushed the claude/cloud-761-issue-key-conformance-fixture branch from 5b06ff5 to fca9130 Compare September 6, 2026 05:11
@sonarqubecloud

sonarqubecloud Bot commented Sep 6, 2026

Copy link
Copy Markdown

❌ The last analysis has failed.

See analysis details on SonarQube Cloud

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant