Skip to content

UK national calibration run-readiness: doctrine, live parity trio, US-format diagnostics, scorer, staging posture (#623) - #743

Merged
juaristi22 merged 30 commits into
mainfrom
uk-national-first-calibrated-candidate-623
Aug 25, 2026
Merged

UK national calibration run-readiness: doctrine, live parity trio, US-format diagnostics, scorer, staging posture (#623)#743
juaristi22 merged 30 commits into
mainfrom
uk-national-first-calibrated-candidate-623

Conversation

@juaristi22

Copy link
Copy Markdown
Collaborator

Part of #623 under #665 (WS-D). Registry #736 item 10 asked to close-or-narrow #623 after #729 delivered the seam core; this PR is the narrowing, per the adjudicated plan (María, 2026-08-21): all calibration tooling lands now — solve doctrine, honest gate evidence, US-format diagnostics, candidate-vs-incumbent scoring, run posture — and the licensed armed run itself stays held until the WS-E spine completes E8 (#684) + E10 (#686), or María names a different base. Builds on merged #735 (the 2025 Ledger surface) and #733 (the FRS 2024-25 line).

What lands

  1. National solve doctrine — new uk_runtime/national_doctrine.py mirroring local_doctrine.py: every solver option the seam previously inherited silently (epochs 256, lr 0.02, max_weight_ratio 10.0, seed 0, loss cap 10.0, l0 0.0, free mass, uniform target weights) is a declared, tamper-tested constant. Doctrine v1 deliberately lifts Add the UK national calibration step over ledger-backed target references #729's values; first-run evidence is the revision path (adjudication 3).
  2. Single compile path + writer-safe materializationUKNationalCalibrationStage consumes the driver's compiled registry at load_uk_frs_release().calibration_year (2025) and materializes bindings through the shared interpreter via a lifecycle subclass of the shared UKFrameTargetAdapter: prepared scratch columns are restored away post-solve, so the staged frame survives the real HDFStore writer (closes Add the UK national calibration step over ledger-backed target references #729 dispositions finding 4, per adjudication 2 — materializer-owned lifecycle). The materialization period is passed explicitly — never the input frame's base-year time_period, which lags the calibration year.
  3. Canonical weight/mass conventionsnational_calibration_mass_reason() names the bound families; a post-solve fence requires the CALIBRATED kind and exactly one appended mass record; the stage manifest carries the rowwise-convention weights block (household_weight_kind_chain, calibration_mass_change, solve block, doctrine echo).
  4. The parity trio evaluates for real on armed builds_stage_parity_evidence builds the uk_export_surface/uk_target_surface/uk_target_fit evidence from independently sourced sides: candidate from the staged frame and solve diagnostics, reference from the frozen eFRS parity instrument (parity_reference parameter, driver passes the committed instrument) and the declared registry at name@period grain. A copied reference can never fabricate a pass; unarmed builds keep the honest evidence_absent, pinned in both postures.
  5. US-format-identical diagnostics + chronicle provenance — the driver's bespoke diagnostics blob is replaced by uk_calibration_diagnostics_payload/write_uk_calibration_diagnostics — the same shared producer the US release path imports (schema v6, per-epoch loss_trajectory, per-target rows, loss attribution), so the calibration dashboard consumes UK and US files identically. A hermetic format-parity pin asserts the UK payload's shared layer equals the shared producer's output with uk_diagnostics strictly additive. The build block carries the chronicle artifact provenance (facts sha256, manifest sha256, profile ids), build id, code pins, and input posture — every target value traceable through ledger_facts_sha256 (Publish Chronicle package IDs in calibration diagnostics targets #661).
  6. Candidate-vs-incumbent scorer (US base v2: one CPS+ACS+PUF-detail pool; datasets labeled by exact record count (dense = full pool; exact-k L0 selection) #578 rule 1) — new tools/score_uk_national_candidate.py: both artifacts rescored on the same frozen register with relative_error_loss (cap 10.0), emitting the June-schema score_vs_enhanced_frs block ({candidate,incumbent}×{train,holdout,full}_loss + target_wins) with per-family wins. Signed differences: holdout_basis: "none_declared" (June's holdout split lives only in the archived pipeline); the diagnostics build block reserves score_vs_enhanced_frs: null until the licensed run merges the real receipt.
  7. Run readiness, runs held — a declared non-certified staging-candidate input posture (--staging-candidate-input-sha256: sha-gated with a mid-read race guard, labeled tier in the build record, refused for release candidates) so the armed run is push-button on any pre-clone spine; docs/uk-national-calibration-runbook-623.md records the exact command, the evidence-dir layout, and the two unblock conditions. Grain basis (adjudication 1): the incumbent's published enhanced_frs_2024_25.h5 is itself pre-clone (52,846 households; clone_and_assign feeds only its local-weights product), so pre-clone scoring is apples-to-apples — with one signed method difference to carry into the score receipt: incumbent national weights are a collapsed local solve, ours is a direct national solve under doctrine.

Explicitly out of scope

Verification

Hermetic: doctrine pinned tests; single-compile-path + writer regression through the real write_uk_national_frame; armed/unarmed parity-trio postures; mass-record fences; US↔UK diagnostics format-parity pin; scorer unit tests on synthetic twins; staging-posture flag matrix (adversarial refusals). Synthetic end-to-end in CI: armed synthetic build → real writer → full battery, US-format diagnostics, Logbook row with ledger_facts bound in input_pins_digest. Full three-shard suite + ruff green locally on the merged tip.

🤖 Generated with Claude Code

juaristi22 and others added 9 commits August 21, 2026 11:39
… re-map, ingest scale ladder

The frozen release object uk/frs_release.json (survey/base 2024, calibration
2025, SN 9563, DOI, UKDS zip sha, HF acquisition pins) drives lockstep asserts
over the re-pinned raw-tab manifest: all 21 frs_table artifacts re-pinned to
the 2024-25 tabs with the SPI-convention keys, six stages' SN 9252 prose
defect fixed, and the typed spec moved in lockstep. TIME_PERIOD and the HMRC
SPI/CGT build periods move to "2024"; the 2023-24 HMRC published surface is
re-mapped as a signed nearest-available-vintage declaration
(period_mapping: latest_published_tax_year) with the frozen original
byte-untouched and the source-contract validator reconstructing the live
canonical payload. take_up_contract build_year 2024 flips the two 2024
date-keyed rates. The ingest driver joins the #627 scale ladder
(--sample-fraction post-frs_spine via the generic frame-sampling helpers)
with receipt-postures on three full-scale fences at sampled rungs, and the
WAS bridge-donor locator defect that refused every full-roster licensed run
is fixed with a hermetic pin-coherence regression test.

Part of #723.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
The regenerator pins move to enhanced_frs_2024_25.h5 at the v1.56.14 tag
(= incumbent-at-pin ebf733c) and the committed reference re-freezes at the
new vintage: 145 columns (surface unchanged), period "2024", entity counts
now test-pinned with the re-derived record-count identity
(16,288 raw + 10,000 SPI) x 2 + 270 CGT band donors = 52,846 - no raw
household is dropped at 2024-25 and the donor count follows the HMRC band
file (30 x 9). The release-input coverage manifest regenerates against the
new reference with the certified 2023 candidate unchanged (candidate side
moves at #686); known gaps stay empty and the restored-column receipts hold.

Part of #723.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
The efrs-post-calibration input-mass descriptor is replaced in place (same
name, #687's replacement model): enhanced_frs_2024_25.h5 identity and the
totals_sha256 of the regenerated 131-column weighted-totals evidence, with
the registry, gates.json, and the data-shard publication mirrors re-pinned
in the same reviewed change per the gate-battery contract. No thresholds
move - #723 records the re-measured baselines, #686 arms them.
UK_REFERENCE_DATASET_NAME follows the incumbent's 2024-25 dataset name. Both
per-reference reviewed exclusions are re-signed against the new reference
(charitable_investment_gifts; owned_land on a fresh 2024-25 stability
receipt), pending the approver's confirmation on the PR.

Part of #723.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
…ise channel-blind scope note

UK_REFERENCE_DATASET_NAME drops the _recalibrated suffix (adjudicated
2026-08-20): the pinned reference is the published enhanced_frs_2024_25.h5
itself and no recalibrated variant exists at this vintage, so the gate-report
label now names the artifact exactly (June report strings keep their own
label). The registry scope notes stop claiming the incumbent "structurally
lacks" the SPI clone channel - the 2024-25 artifact carries the synthetic
rows structurally but no admin-restored mass in the channel-exclusive
columns, which is the fact the reviewed exclusions rest on; the approved
exclusion reasons are untouched. Gate digests re-cut over the post-#729
union in the same change.

Part of #723.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
…block; self-describing June freeze

Finding 1 (accepted): the family-coverage block rendered the #723 re-mapped
period fields from the canonical manifest while hashing only the frozen
mirror - evidence fields and their hash must name the same bytes. The block
now carries a dual pin (source_manifest for the frozen June identity,
canonical_source_manifest for the bytes the re-mapped fields come from) and
a test binds each field set to the sha256 of the file it actually derives
from. Finding 2 (rejected with armor): the committed replay report is the
June evidence freeze and deliberately keeps mapped_build_period 2023 - it is
evidence for the grandfathered release, not the 2024 line, and retires with
the frozen manifest after #686 per #687; instead of regenerating it, a new
assertion binds it to the FROZEN manifest's declared mapping so the
partition is self-describing.

Part of #723.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
…tree

Both parents edit uk/gates.json and uk/country_package.json, so the merged
manifest digests differ from either side's pins. Re-pinned by recomputation
(never by picking a side): spec bundle e12a2cb8…, policy 404968fb…,
gates manifest 59c7808d…, spec fingerprint bfb98736….

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
…-format diagnostics, scorer, staging posture

Work items 1-7 of the narrowed #623 plan (runs held per the 2026-08-21
adjudications; everything here is synthetic/hermetic):

- uk_runtime/national_doctrine.py: every solver constant the national
  calibration inherits becomes a declared, tamper-tested doctrine value
  (epochs 256, lr 0.02, ratio 10.0, seed 0, loss cap 10.0, l0 0.0,
  free mass, uniform weights).
- Single compile path: the stage consumes the driver's compiled registry
  at the release calibration year and materializes bindings through the
  shared target_materialization interpreter; prepared scratch columns are
  restored post-solve, so the staged frame survives the real HDFStore
  writer (closes #729 dispositions finding 4).
- Canonical mass/weight conventions mirroring the rowwise path: declared
  mass reason, post-solve fence (CALIBRATED kind, exactly one appended
  record), household_weight_kind_chain + calibration_mass_change manifest.
- The parity trio evaluates for real on armed builds: candidate side from
  the staged frame and solve diagnostics, reference side from the frozen
  eFRS parity instrument and the declared registry at name@period grain —
  never a copied reference; unarmed builds keep evidence_absent.
- calibration_diagnostics.json is now produced by the same shared
  diagnostics producer the US release path uses (per-epoch loss
  trajectory, per-target rows), with a pinned US<->UK format-parity test;
  the build block carries the chronicle artifact provenance
  (facts/manifest shas, profile ids) and reserves score_vs_enhanced_frs.
- tools/score_uk_national_candidate.py: #578 rule-1 scoring, both
  artifacts rescored on one frozen register, June-schema score block with
  a declared none_declared holdout basis.
- Declared non-certified staging-candidate input posture (sha-gated with
  a mid-read race guard, refused for release candidates) and the held-run
  runbook, unblocking on WS-E E8+E10 or an explicitly named base.

Implemented by Codex from the reviewed plan; review pass fixed the
materialization period (declared calibration year, never the frame's
base-year time_period) and the parity-evidence sourcing above.

Part of #623 under #665.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
…d adapter

Both PRs merged upstream, including María's adaptations of parts of this
branch's work (c534517 routes calibration through the shared materializer;
1d066e5 is the cherry-picked #733 review fix). Reconciliation:

- target_materialization.py and verify_uk_identity_stability.py: main's
  reviewed versions taken wholesale.
- national_calibration.py: this branch's registry+period+doctrine stage kept
  (all tests and the driver target it); its private frame adapter replaced by
  a lifecycle subclass of main's shared UKFrameTargetAdapter — one adapter,
  now writer-safe (prepared scratch columns restored away post-solve, per the
  adjudicated materializer-owned lifecycle).
- Main's packaged-binding stage tests ported to the registry+period API; the
  persistence assertion inverted to the adjudicated writer-clean invariant
  (result columns exactly equal the input columns); the materialization stub
  adapter adopts the shared count-variable convention.
- Digest pins auto-merged to main's post-union re-cut and verified green by
  the pin suites — no re-measurement needed.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
…ain merge

The post-#735/#733 merge at 9a8bd46 resurrected the standalone
`except ValueError` block that #733's review commit 1d066e5 had
folded into the single generic handler. Because the narrower clause
catches first, the named-edge branch inside `except Exception` became
unreachable, silently reverting a reviewed structural fix.

This branch has no business touching the spine driver at all — its
scope is the national calibration seam — so the file returns to main's
version exactly.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
juaristi22 added a commit that referenced this pull request Aug 23, 2026
An armed UK national build crashed before its first stage: _source_pins
stored the full Ledger provenance block under the 'ledger_facts' role, but
role_pins_digest requires each role to carry exactly sha256 and size_bytes,
so every armed run raised ValueError at input_pins_digest.

The two halves landed independently — the strict pin contract with Logbook
adoption (#666) and the Ledger role pin with the compile-parity wiring — and
neither PR branch fires it alone, because only an armed run (--ledger-facts)
reaches this path and the licensed runs were held. It reproduces on both
#743 and #747 branches and on main.

The pin now carries the feed's verified digest and byte size; the richer
Ledger identity block already travels in safe_artifacts, source_vintages,
and the diagnostics build block, so nothing is lost.

This fix belongs upstream in #743 (the run-readiness PR whose runbook
documents the armed command); it rides the build branch until then.
juaristi22 added a commit that referenced this pull request Aug 23, 2026
… production path

First armed run receipt: 187 of 388 references skipped at materialization,
because they bind simulated tax-benefit outputs (income tax, NICs, UC,
caseloads, payment-band crosstabs) and the calibration stage's frame adapter
reads stored columns only. UKPolicyEngineAdapter — the live-sim adapter —
has no production caller, and the stage's tests feed synthetic frames with
measure columns pre-attached, so every real input (spine or certified June
candidate) fails the 388-target surface the same way. The June release never
hit this because its 149 targets were demographics-only.

The runner now drives the materializer's own skip reports: each round
computes the missing (entity, variable) pairs from a live policyengine-uk
simulation over the same records — native-entity calculate, numeric map_to,
categorical group-to-member broadcast, boolean any-collapse — attaches them
to the frame, and retries until the register binds. The incumbent gets the
same treatment before scoring, so rule 1 compares both datasets under one
yardstick, and the score block declares that symmetry. References this
posture cannot bind (the salary-sacrifice counterfactual deltas need
adapter.counterfactual_delta) are excluded with a per-reference receipt.

The production fix belongs in the #743 lane: either wire
UKPolicyEngineAdapter into the stage or land this materialization as a
declared pre-calibration step.
juaristi22 and others added 8 commits August 24, 2026 12:58
An armed UK national build crashed before its first stage: _source_pins
stored the full Ledger provenance block under the 'ledger_facts' role, but
role_pins_digest requires each role to carry exactly sha256 and size_bytes,
so every armed run raised ValueError at input_pins_digest.

The two halves landed independently — the strict pin contract with Logbook
adoption (#666) and the Ledger role pin with the compile-parity wiring — and
neither PR branch fires it alone, because only an armed run (--ledger-facts)
reaches this path and the licensed runs were held. It reproduces on both
#743 and #747 branches and on main.

The pin now carries the feed's verified digest and byte size; the richer
Ledger identity block already travels in safe_artifacts, source_vintages,
and the diagnostics build block, so nothing is lost.

This fix belongs upstream in #743 (the run-readiness PR whose runbook
documents the armed command); it rides the build branch until then.
243 of the 388 UK references are banded, and none of them were ever
sliced. Every employment-income band materialized 35,351,186 (roughly
everyone with employment income) and every state-pension band 13,518,447
(roughly every state pensioner), so a band's apparent overshoot -- up to
6,759x on the state-pension 1m+ cell -- was an unsliced total compared
against a band value, not a data defect.

The slice was declared but unconsumed: the contract binding names
groupby_variable, each compiled spec carries its own band's lower edge in
Ledger filter metadata, and _prepared_column_values read neither. The
binding's own filters list did work (verified: a family_type == SINGLE
binding correctly masks a COUPLE row), so this is specifically the band
dimension.

Both published encodings reduce to one lower edge -- a numeric
*_lower_bound (HMRC SPI) or a range label in monthly units scaled by
band_period_factor (DWP awards) -- because no reference anywhere declares
an upper bound. A band's upper edge is its sibling's lower edge within the
same contract target, grouped per contract target rather than per
dimension so two measures sharing a dimension cannot slice each other on
the wrong boundaries; the top band runs to infinity. Validated against the
real register: 243 banded references, zero unreadable, and the derived
bounds match the measure names (150000-200000, 1000000-inf, and UC monthly
500.01-600.00 to annual 6000.12-7200.12).

A band whose edge cannot be read now raises, so the measure is skipped and
reported rather than silently reporting the whole population as one band --
the failure mode that hid this. Tests cover the partition property, an
adjacent-bands-differ regression, the monthly label conversion, and the
refusal.

Lives in the shared materialization module, so it cherry-picks into any
branch.
All 15 two-child-limit references failed to materialize. Two naming faults,
both in the contract rather than the provider:

- eight children references declared value_variable "children_count", which
  is not a policyengine-uk variable. They now bind what they actually count:
  uc_is_child_limit_affected for the affected-children rows (mapped to
  household it sums to the count of flagged children) and is_child for the
  children-in-affected-households rows (the total children there).

- fifteen bindings put a prose label in count_of ("affected_households",
  "affected_children", "children_in_affected_households"). count_of is a
  column fallback consulted when value_variable is an entity-count
  indicator, so the provider looked those labels up as columns and raised.
  The labels move to notes, leaving household references on the
  household_count unit indicator as intended.

The provider keeps its documented behaviour; the pre-existing test covering
count_of as a real column still passes.
The stub adapter and national-stage frame both modelled the old
children_count column. Mapped to household, uc_is_child_limit_affected sums
to the number of flagged children, so it serves as both the affected flag
and the affected-children count; the fixtures now carry counts rather than
indicators.

Caught by running the suites properly: the earlier chain piped pytest into
tail, so its exit status was tail's and two real failures rode through into
760fe8b.
UKFrameTargetAdapter.household_condition built a group-membership column
for every non-household entity, so a condition declaring entity "person"
asked for "person_person_id" and raised. People sit directly in a
household; only group entities need that lookup.

Every ONS household-composition reference uses person-level conditions
(is_child, age), so all ten were excluded from calibration. That left
household structure unconstrained while population stayed targeted, and
the solve satisfied population by inflating households rather than
multiplying them: weighted benunits per household reached 1.476 against
the incumbent's 1.132, from an identical unweighted 1.158.

The downstream cost is Universal Credit. 75.5% of the candidate's single
UC claimants end up in multi-benunit households (incumbent: 8.6%), where
only one benunit claims housing costs, so the housing element reaches
47.2% of them against the incumbent's 73.9% -- despite the candidate
having more renters (88.3% vs 79.2%) and higher rent, and despite beating
the incumbent within both strata (92.6% vs 78.9% solo, 32.5% vs 20.4%
multi). Simpson's paradox, driven entirely by composition. Median single
UC award falls to GBP 8,059 against GBP 9,310, emptying the GBP 8-11k and
GBP 14-19k award bands that carry 98% of the caseload shortfall.

Shared UK runtime, so this cherry-picks into the spine lane.
The 1a3274b crash fix landed without a test; this locks the contract it
restored: the ledger_facts role pin carries exactly {sha256, size_bytes}
(both feed layouts) and survives role_pins_digest, while the full
provenance block is rejected — the exact shape that crashed the first
armed run at input_pins_digest.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
…ine rule into the solve

Re-cut of the build branch's 035ed63 under María's 2026-08-24 ruling:
family_equal enters _ALLOWED_TARGET_WEIGHT_RULES with its armed-run
receipts, and uk_national_target_loss_weights() derives the family-share
vector — but the doctrine DEFAULT stays uniform. She passes family_equal
as an explicit per-run setup while the weighting doctrine is measured;
neither rule is adopted by default (run-9 receipts cut both ways).

Also fixes the latent wiring defect the armed run exposed: the stage
echoed target_weight_rule in its manifest but never passed a weight
vector to calibrate(), so any declared rule silently solved uniform.
The vector now travels explicitly; uniform maps to None, keeping the
shipped identity byte-stable under the default.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
The two-child-limit rebinding (760fe8b) replaced the prose count_of
labels with real value_variable columns; the chronicle loader-guarantee
test still demanded count_of on every baseline_flag_crosstab binding.
The guarantee now matches the provider: affected_flag_variable plus
either counted-column spelling. Latent on the build branch too — this
suite was never re-run there after the rebinding.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
@vahid-ahmadi

Copy link
Copy Markdown
Contributor

Automated review pass (Claude Code, high effort, diff only — no verification runs). Posting on the draft since two of these bear on whether the calibration gate means anything; the other two are dead code.

Substantive

1. packages/microcosm-build/src/microcosm/build/uk_runtime/hmrc_restoration.py:851 — the certified-pin check is skipped for the staging tier. _validate_certified_candidate_identity returns early for any identity whose tier == "staging_candidate", verifying only that revision equals a hard-coded constant; the sha256/size comparison against the certified pin never runs. Any UKCertifiedCandidateIdentity constructed with that tier — not only one produced by verify_staging_candidate_uk_input — passes the HMRC replay-base gate with an unverified file. The SHA binding currently lives only in the CLI path, which makes it a convention rather than an invariant; a second caller constructing the identity directly gets no binding at all.

2. tools/score_uk_national_candidate.py:60 — the rule-1 holdout comparison is vacuous by construction. The score block sets candidate_train_loss, candidate_holdout_loss and candidate_full_loss all to the same candidate.final_loss (and likewise for the incumbent), with holdout_basis: "none_declared". Since the candidate was calibrated on this very registry, there is no held-out quantity in the comparison and rule 1 can only ever be won by the candidate — any overfit divergence is silently absorbed. This is the same failure shape as the E7 receipt on #747: a green result that reflects the absence of a check rather than the presence of agreement. Given this is the component deciding whether a calibration run succeeded, it seems worth either declaring a real holdout basis or having the rule refuse when holdout_basis == "none_declared" rather than scoring.

Dead code

3. tools/build_uk_frs_spine.py:871 — the new except ValueError as error: clause is inserted ahead of the pre-existing except Exception as error: whose body branches on isinstance(error, ValueError). That branch is now unreachable, and any ValueError special-casing it performed beyond the rung-abort signature is silently lost. Worth a look given this is the same handler restructured for the #733 review — it may be that the merge of the two paths dropped something.

4. tools/build_uk_national_dataset.py:1202source_vintages["frs"] = load_uk_frs_release().vintage is added two lines above an identical pre-existing assignment; the new line is dead and load_uk_frs_release() is now called twice.

Findings 3 and 4 are trivial. 1 and 2 are the ones I would want resolved before this leaves draft — in both cases a gate reports success without having tested the thing it names.

juaristi22 and others added 3 commits August 24, 2026 15:37
…d measure-resolution loop

Two moves the calibration seam needs, both behaviour-neutral for the June
path. The H5 reader/writer pair moves next to uk_national_frame and
validate_uk_national_frame — its natural home — with national_build
re-exporting every name, so the legacy driver and its suites are
unchanged. The seam can then load and stage frames without importing the
legacy build module at all.

target_materialization grows the country-agnostic half of the assessment
harness's resolution loop: probe-adapter rounds, once-per-round dedup of
a shared missing key, and a fail-loud finish. The country-specific step
is injected as a provider, so nothing about policyengine-uk enters the
shared module. Where the harness excluded an unresolvable reference with
a receipt, production raises: register pruning belongs to a reviewed
exclusion register, never to a silent loop.

assert_calibration_input_finite is the seam's NaN fence — it lists every
offending column at once instead of failing on the first.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
…overrides

187 of the 388 compiled references bind model outputs — income tax, UC —
that no production code ever computed onto a frame; the stage's adapter
reads stored columns only, which is why the first armed run could not
bind them. The provider computes them through policyengine-uk on four
receipted routes and injects them at adapter level, where names are
table-scoped, so the Frame's global column-uniqueness rule is never
touched and the staged artifact carries no scratch column.

The five salary-sacrifice counterfactual references cannot bind without
a live counterfactual run, so they leave the register through a reviewed
exclusion file with a written reason each, not through a silent skip. A
stale entry that matches nothing is an error.

Doctrine defaults stay the reviewed constants. uk_doctrine_with_overrides
builds an effective doctrine through the frozen dataclass, so the closed
vocabularies still apply, and returns the diff against v1 — recorded as
an explicit deviation wherever the run is described. Only the four fields
calibration work actually turns are overridable; seed, weight ratio,
mass and scale rules stay reviewed constants and refuse.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
…re given

The June national driver rederives SPI income, redraws capital gains and
replays support tables onto its input — work the WS-E spine now ships
already derived, and which its own zero-weight precondition refuses to
perform on a spine at all. This seam does none of it: it verifies the
input pin, loads, fences NaNs, solves weights under doctrine, and writes
the staged H5, US-format diagnostics, build record and signed gate
report. Its defining invariant has a test: every data column of the
staged artifact is identical to the input, and only the household
weights, their kind and one appended mass record differ.

Nothing here imports the legacy build module, so the June path can retire
on its own schedule without this lane moving.

The gate posture is calibration-scoped. Spine-construction gates cannot
pass a build that constructs no spine, and pretending otherwise would be
the invented pass this repo refuses; each out-of-scope entry is listed in
the signed report with its reason, and a test asserts scope and
exclusions together cover every declared gate, so a new gate cannot land
unclassified. Aggregate-admin evidence follows the anchors' own
convention — per-household means for the NEED anchors, a national total
for NHS — and refuses when the frame cannot supply an anchor rather than
reporting a zero.

The build record states shippable: false with its reason. A publishable
artifact needs the full battery, which is release-cut work.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
…ntage line

Finding 2 (confirmed). The score block copied final_loss into
candidate_holdout_loss and incumbent_holdout_loss while declaring
holdout_basis none_declared. The June fixture this block mirrors carries
a genuinely different holdout value — 0.1239 against a 0.0159 train loss,
the eightfold degradation that made its holdout the discriminating
comparison — so repeating the fitted loss there reads downstream as
perfect generalization from a measurement never made. The keys now report
absence, and declaring a basis without computing its split raises rather
than falling back. The scorer also now takes the H5 loader from its new
home instead of the legacy build module.

Finding 4 (confirmed). source_vintages["frs"] was assigned twice, so
load_uk_frs_release() ran twice; main carries one line, which makes this
a branch-local artifact of the same merge that duplicated the spine
driver's failure handler. A sweep of tools/ and uk_runtime/ for the same
defect class found no others.

Finding 1 (refuted, comment only). The staging-tier early return skips
the certified-pin equality, which is what the tier declares, but the
bytes are bound before it: the verification token is a module-private
sentinel only the two verifiers stamp, and the staging verifier hashes
the file against a mandatory declared sha with a mid-read race guard. The
comment now says so, since the shape invites the reading twice now.

Finding 3 was already fixed on this branch before the review was posted.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
@juaristi22

Copy link
Copy Markdown
Collaborator Author

Thanks — dispositions below. Two confirmed and fixed, one refuted with evidence, one that was already fixed on the branch before the review landed.

A timing note first: the review read origin's then-head 9a8bd46d. ccb80a3d was still local at that point, and the branch has since grown the calibration seam itself (María's ruling: everything calibration folds into this PR, and the seam must not depend on the legacy national build).

2 — scorer holdout comparison: CONFIRMED, fixed

You're right that the block reports a holdout it never measured, and the June fixture makes it concrete. packages/microcosm-data/tests/fixtures/uk_june_2023/.../calibration_diagnostics.json carries genuinely distinct values:

train holdout full
candidate 0.0159 0.1239 0.0376
incumbent 0.0385 0.3784 0.1069

June held rows out, and its holdout was the discriminating comparison — an eightfold train-to-holdout degradation on the candidate. Copying final_loss into that key reproduces June's shape with none of its meaning: a consumer reading candidate_holdout_loss sees train and holdout agreeing exactly, which reads as perfect generalization from a measurement never made. The sibling declarations (holdout_basis: "none_declared", loss.train_equals_full: true) were honest but don't help a consumer reading the key itself.

Fixed by reporting absence instead of a fabricated number: both holdout keys are now null while holdout_basis is none_declared, and declaring a basis without computing its split raises NotImplementedError rather than falling back to the fitted loss. Same shape as the evidence_absent-over-invented-pass rule the battery already applies, and as the parity-trio un-aliasing earlier in this lane. Tests now pin the absence; the assertion that previously pinned candidate_holdout_loss == candidate_full_loss was itself locking in the defect.

Two refinements to the finding. "Can only ever be won by the candidate" is stronger than what holds — the scorer's own synthetic-twins test has the candidate losing on full loss (0.25 against the incumbent's 0.1), and the candidate's structural advantage is declared, not hidden, in signed_asymmetries.incumbent_own_registry. And I did not take the second remedy — refusing to score when the basis is none_declared — because #578 rule 1 as adjudicated is a fit comparison of both artifacts rescored on one frozen register, not a generalization test; refusing would block the rule entirely rather than repair it. Declaring a real holdout basis is genuinely worth doing and is a separate piece of work.

4 — duplicated source_vintages["frs"]: CONFIRMED, fixed

Confirmed, and worth recording why it was there: main carries one such line, this branch carried two, so it is a branch-local artifact of the same 842e263b merge that resurrected the spine driver's duplicate failure handler in finding 3. One merge, two duplicate-code artifacts; María caught one at ccb80a3d, this one survived.

Since that's now a defect class rather than an incident, I swept every file in tools/ and uk_runtime/ for repeated identical assignments in close proximity. These two are the only real instances in the UK lane — the remaining hits are legitimate parallel branches, e.g. tools/build_uk_frs_spine.py:405 and :407, which are distinct elif operation.kind arms that happen to assign the same thing.

1 — staging-tier certified-pin skip: REFUTED, comment upgraded

The load-bearing claims don't hold against the code. Before the tier branch is reached, _validate_certified_candidate_identity refuses any identity whose _verification_token is not the module-private sentinel built at hmrc_restoration.py:97, and then refuses any identity lacking a verified _source_file_fingerprint — both checks apply to the staging tier. Exactly two functions stamp that token, and the staging one binds the bytes itself:

def verify_staging_candidate_uk_input(path, *, expected_sha256) -> UKCertifiedCandidateIdentity:

It takes the digest as a mandatory keyword, hashes the file, re-reads the source fingerprint to catch a mid-read swap, and raises on mismatch. So the SHA binding is not "only in the CLI path" — it is inside the verifier — and a caller cannot construct a passing identity without reaching for a private sentinel.

What the early return does skip is the certified-pin equality, which is the tier's entire purpose: a declared, non-certified input, labelled non_certified_staging_candidate in the build record, with --staging-candidate-input-sha256 refused alongside --release-candidate. The fair residual is that the digest is operator-declared, so it binds "the file the operator named" rather than a certified artifact — which is what the tier declares, and it is recorded.

No behaviour change, but the shape has now invited this reading twice, so the early return carries a comment naming the two guards that do the binding. Worth noting this is legacy June-path code that the new seam does not touch at all; its retirement sits in #757.

3 — spine driver's unreachable handler: already fixed

ccb80a3d ("Drop the duplicate spine-driver failure handler reintroduced by the main merge") removed the standalone except ValueError block; the file has a single handler whose isinstance branch is reachable, and nothing beyond the rung-abort signature was lost. That commit was local when the review was posted and is now on the branch — same merge artifact as finding 4, as above.

@juaristi22

Copy link
Copy Markdown
Collaborator Author

Scope update: the calibration seam now lands here, with the run campaign's receipts

This PR opened as run-readiness tooling with the licensed run held. Since then the first armed campaign actually ran (9 runs on the uk-first-calibrated-run-623-686 build branch), and it changed what "ready" means: it surfaced defects that were invisible while runs were held, and it proved the production driver cannot calibrate a WS-E spine at all. Per María's rulings (2026-08-24/25) the fixes and the new seam fold into this PR rather than a stack of follow-ups. Commit-by-commit map below, each with the evidence that motivated it.

Campaign fixes (upstreamed from the build branch)

Commit What Run evidence
43d3e2a5 Ledger role pin carries exactly {sha256, size_bytes} Every armed build crashed at input_pins_digest before its first stage. Reproduces on both PR branches and on main; only --ledger-facts reaches the path, which is why held runs never saw it
53a66094 Regression test for that pin shape The fix arrived without one; the test pins both feed layouts and asserts the full provenance block is still rejected
8eef7eb4 Banded groupby measures are actually sliced 243 of 388 references returned the unsliced population total — every employment-income band 35,351,186, every state-pension band 13,518,447. Fixing it moved initial loss 5.227 → 0.686, targets-within-10% 20/251 → 295/351, and the weight-ratio gate from failing to passing
f8ecfa71 + f24b401b Two-child-limit references bound to real variables All 15 were dead: 8 bound the non-existent children_count, 15 put a prose label in count_of. After: households_affected +0.3%, children_affected +1.4%, from previously unmaterializable
aa2fd22a Person-level household conditions reduce through person_household_id The reducer built person_person_id for any non-household entity, so all ten ONS composition references raised — household structure was simply unconstrained (benunits/household 1.476 vs the incumbent's 1.132)
91074ddd Loader-guarantee test follows the corrected crosstab shape Latent on the build branch too: that suite was never re-run there after the rebinding
e7d392be family_equal enters the target-weight vocabulary; doctrine rule reaches the solver Two things. The vocabulary is motivated (under uniform, the hmrc family is 57% of the surface and supplied 101 of 102 past-cap references). Separately, a latent production defect: the stage echoed target_weight_rule in its manifest but never passed a weight vector to calibrate(), so any declared rule silently solved uniform. Per María's ruling the default stays uniform — 1500 epochs and family_equal are per-run setups she passes, not defaults

Deliberately not picked: the epochs-1500 default (same ruling), and age_tail — it writes person.age, so it belongs to the spine lane, where it has since landed in #747 as a declarative source stage.

The calibration seam (new scope)

The production driver rederives SPI income, redraws capital gains and replays support tables onto its input — work the WS-E spine now ships already derived, and which its own zero-weight precondition refuses to perform on a spine at all. So a spine-based calibration was not a matter of flags; there was no code path. Per María's architecture ruling the seam does not wrap national_build.py: country-agnostic contributions extend the existing shared modules, UK-specific code lives in new UK files, and nothing imports the legacy build.

  • 96e73047 — moves the UK national H5 reader/writer next to uk_national_frame/validate_uk_national_frame (with national_build re-exporting, so the June path and its suites are untouched), and adds the country-agnostic measure-resolution loop plus the NaN fence to target_materialization.py. The loop is the harness's fixed-point rounds with one production change: where the harness excluded an unresolvable reference with a receipt, production raises — pruning belongs to a reviewed register, not a silent loop.
  • 971a9699 — the UK simulated-measure provider. 187 of 388 references bind model outputs (income tax, UC) that no production code ever computed onto a frame; the stage's adapter reads stored columns only, which is exactly why the first armed run could not bind them. Values are injected at adapter level, where names are table-scoped, so the staged artifact carries no scratch column. The five salary-sacrifice counterfactual references leave the register through a reviewed exclusion file with a written reason each. Also uk_doctrine_with_overrides: per-run overrides validated through the frozen dataclass and recorded as an explicit diff against v1, with the four campaign knobs overridable and seed/ratio/mass/scale rules staying reviewed constants.
  • e54e546b — the seam and its driver. Verify the pin, load, fence NaNs, solve under doctrine, write the staged H5, US-format diagnostics, build record and signed gate report. Its defining invariant has a test (test_seam_never_modifies_data_variables): every data column of the staged artifact identical to the input, only household weights, their kind and one appended mass record differ. Gate posture is calibration-scoped — spine-construction gates cannot pass a build that constructs no spine, and claiming otherwise would be the invented pass this repo refuses, so each out-of-scope entry is listed in the signed report with its reason and a test asserts scope ∪ exclusions covers every declared gate. Build records state shippable: false with the reason: a publishable artifact needs the full battery, which is release-cut work.
  • 38bb5f74 — Vahid's review dispositions (separate comment above).

Verification

Full per-shard suite green on the pushed tip — frame, fit, calibrate, build and data all exit 0 — plus ruff check . clean. Pin suites specifically re-run to prove no digest moved (test_spec_engine_country_bundles, test_gate_battery_contract_pins, the microcosm-data contract tests). The June driver's own suites are part of that and are unchanged.

Explicitly not here

Registry coverage: 29 contract targets carry no reference row and so never compile. That is verified and written up as an adjudication package in #736 — it needs several rulings and a digest re-cut, so it does not ride this PR. The first production calibration run is likewise a separate step, once this and #747 land.

🤖 Generated with Claude Code

juaristi22 and others added 2 commits August 24, 2026 18:22
…ibrated-candidate-623

# Conflicts:
#	packages/microcosm-build/src/microcosm/build/uk/country_package.json
#	packages/microcosm-build/tests/test_country_spec.py
#	packages/microcosm-build/tests/test_spec_engine_country_bundles.py
The wheels gate installs the base wheels without the optional pandas HDF
backend, so the three calibration-run tests that write and re-read a real
staged H5 failed there with an ImportError while passing everywhere else.
They now carry the repo's existing guard, matching every other test that
touches the HDF backend.

The two other new seam test files monkeypatch the writer rather than
performing real H5 I/O, so they need no guard and already passed the gate.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
@juaristi22
juaristi22 marked this pull request as ready for review August 25, 2026 07:47
@vahid-ahmadi

Copy link
Copy Markdown
Contributor

Second pass (Claude Code, high effort, diff only — no build or test execution), scoped to the new material: the calibration seam (96e73047, 971a9699, e54e546b) and the upstreamed campaign fixes.

First, on the earlier round: the refutation of finding 1 is accepted and it was my error. The reviewer read the early return without tracing the guards above it — _verification_token against the module-private sentinel and the verified _source_file_fingerprint both apply before the tier branch, and verify_staging_candidate_uk_input takes expected_sha256 as a mandatory keyword and hashes the file itself, so the binding is in the verifier rather than only in the CLI. Your residual framing — that the digest is operator-declared, and so binds the file the operator named rather than a certified artifact — is the accurate statement. The clarifying comment is a good outcome for a shape that has now misled two readers.

Four findings on the new scope.

1. uk_runtime/calibration_run.py:~430 (_aggregate_admin_totals) — clones collapse, and the anchor lands at half mass

Person-anchor weights are built with dict(zip(household_id, household_weights, strict=True)). strict=True checks that the two iterables are the same length; it does not check key uniqueness. Since cgt_incidence_clone duplicates every household with the mass split 0.5, each household_id appears twice, and the dict keeps only the last clone's weight. Every person in that household is then weighted by one clone rather than the pair, so the admin anchor is measured at roughly half its intended mass.

The failure is silent and it looks like a calibration problem rather than a bug — the anchors would simply sit low against their targets, which is exactly the kind of signal that gets chased into the solver. Worth either summing weights per household_id explicitly or asserting uniqueness and refusing, so a cloned input can't be measured this way by accident.

2. target_materialization.py:~530 (_band_lower_edge) — the band edge isn't tied to the groupby variable

The edge is taken from the first alphabetically-sorted ledger_filter_*_lower_bound key (or the first range-label-matching ledger_filter_*), with no check that the key relates to binding["groupby_variable"]. A spec carrying two ledger filters with numeric or range metadata — an age filter alongside an income band, say — slices the measure on the wrong variable's edges and silently produces a wrong subpopulation total.

That is the same class of defect 8eef7eb4 set out to end, one layer up: the 243-of-388 unsliced-total bug was loud once measured because the numbers were identical across bands, but a wrong-variable slice produces plausible-looking distinct numbers instead. Binding the lookup to the groupby variable's own filter key would close it.

3. uk_runtime/measure_simulation.py:~150 (UKMeasureResolver.knows) — the fence is unreachable

return native == entity or entity in _ENTITY_ID is vacuously true for any of person / benunit / household, so knows never returns False for a variable that exists. The provider does not know {entity}.{variable} fence in resolve_target_measures can therefore never fire, and an unmappable categorical surfaces later as a wrapped provider failed computing … exception instead of the clean refusal the fence was written to give. Given the new resolution loop raises where the harness excluded-with-a-receipt, the quality of that refusal matters more than it used to.

4. target_materialization.py:~285 (_inject_measure_inputs) — the no-mutation claim needs confirming

The docstring states "The source frame is never mutated", but injection writes adapter.tables[entity][variable] = values directly. That holds only if UKFrameTargetAdapter.__init__ deep-copies the frame's DataFrames, which isn't visible in the diff — the only copying I can see is the test stub and _original_tables. If it aliases them, every probe round appends scratch columns to the caller's input frame, and the "no scratch column in the staged artifact" property rests on restore() alone rather than on the frame being untouched.

Not necessarily a defect — but it is the invariant test_seam_never_modifies_data_variables exists to protect, and a test that runs after restore() would pass either way. Worth confirming which of the two is true, and if it is aliasing, either deep-copying at construction or restating the docstring to claim what actually holds.


Finding 1 is the one I would hold the merge for. 2 is close behind — both produce wrong numbers without producing an error, which is the failure mode this lane has been systematically closing.

On the scope change: folding the campaign fixes and the seam in here rather than stacking follow-ups seems right given they are what "ready" turned out to mean, and the run evidence in your table makes the case for each one well. The banded-measure fix moving initial loss 5.227 → 0.686 and targets-within-10% from 20/251 to 295/351 is a striking result to have surfaced only once runs were armed.

@MaxGhenis

Copy link
Copy Markdown
Contributor

Ran a deep automated review of this PR through my agent session before your merge (sol lane, adversarial-probe style; CI green is acknowledged — these are checks below the CI waterline). Verdict from the run: hold for five findings, each reproduced by a committed probe rather than asserted. Full code-cited report: REVIEW-743.md on review/pr-743-audit (reviewed head 74b4d768, which matches the current PR head).

The five, ranked:

  1. The documented held-run command omits the required resolver/exclusions arguments and is expected to abort as written.
  2. The rule-1 scorer cannot score exported production measures and does not authenticate the incumbent artifact.
  3. Eight child-count bindings collapse booleans with any(...), yielding household indicator values where child totals are declared.
  4. Release-candidate membership is operator-selectable through exclusions that lack the owner/approval/expiry receipt fields the register schema requires elsewhere.
  5. The calibration pipeline derives an undeclared Logbook scope and does not record unsuccessful attempts.

Your call entirely — if any of these misread the code, say so and merge; the probes and their outputs are in the report so each claim is checkable in minutes. Happy to pair on fixes or have my agents draft them if useful. (#3 seems the most substantive: it would bias the UC family-composition targets the #735 surface just activated.)

juaristi22 and others added 5 commits August 25, 2026 13:54
…ibrated-candidate-623

# Conflicts:
#	packages/microcosm-build/tests/test_spec_engine_country_bundles.py
…band

Three review findings on the calibration seam's measure surface, each one a
wrong number that no error would have surfaced.

The two-child-limit contract declares eight rows whose value is a count of
children over a boolean member variable. The crosstab read that boolean at
the target's own grain, so the resolver's person-to-household route collapsed
it with `any` and published a household indicator against a child-count
target. Counts now declare `value_reduction`, the adapter grows the numeric
sibling of `household_condition` to sum members, and an adapter that cannot
reduce refuses instead of falling back to the same-grain read. The stub and
fixture frames that encoded the old assumption in a comment now carry real
person rows, so the distinction is exercised rather than asserted.

Band edges were taken from the alphabetically-first band-like Ledger filter
with no tie to the binding's groupby variable, so a spec carrying two banded
dimensions could slice on the wrong one — distinct, plausible, wrong numbers,
the same class as the unsliced-measure defect one layer up. The edge is now
attributed to the declared banding dimension, `band_filter_dimension` names it
where the publisher's and the model's spellings differ, and an unresolvable
tie raises.

`UKMeasureResolver.knows` was vacuously true for every entity, so the
resolution loop's refusal could never fire and an unmappable categorical
surfaced as a wrapped provider exception. It now answers for the routes
`compute_uk_measure_input` actually takes.

Reported by @vahid-ahmadi (2, 3) and @MaxGhenis (1).

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
…hat it verified

The seam declared pipeline `uk-national-calibration`, which derives the scope
`uk/national` — not in the ratified vocabulary. The FRS line's spine, staging,
imputation and calibration stages share one chain by design (the dataset token
names the base data, not the build mechanism), so calibration is
`uk-frs-calibration` and lands on `uk/frs`.

Three more lifecycle gaps went with it. The predecessor digest was resolved at
the very end, so a disagreeing chain refused only after a staged H5,
diagnostics and a signed gate report already existed; it is now validated
before anything is written. Only successful runs recorded a row, against the
binding rule that successful, failed, refused and discarded attempts all
produce one; the attempt is now wrapped, and a refusal writes an error receipt
and a `failed` row while staging nothing. And the build id was deterministic
per release id, which both the local chain and the store reject on a re-run;
attempts now carry unique ids and the release id travels in run_config.

`run_uk_calibration` also accepted `ledger_artifact` and never used it, so the
facts and manifest digests the caller had just verified reached neither the
identity digest, the build record, nor the Logbook row. The verified identity
is now sealed into run_config; a bare feed's missing manifest is recorded as
absent rather than invented.

Release candidates no longer accept `--measure-exclusions`: the register prunes
the target surface before the solve and the calibration-scoped battery carries
no target-surface gate, so an operator file could narrow what a candidate was
measured against unnoticed. The applied receipt now keeps each exclusion's
tracking reference, and the loader requires one.

Reported by @MaxGhenis (findings 4, 5, 8).

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
…oves

The scorer loaded both H5s and called `score_targets` directly. Every packaged
UK reference binds a slash-named prepared measure, and calibration deliberately
strips those before export — so on a production register it bound nothing and
aborted, while its tests stayed green on plain `measure_a` columns no register
uses. Both sides now go through the same resolve-inject-materialize route the
calibration seam takes, via `prepare_uk_target_frame`, and a skipped target
refuses rather than quietly shrinking the surface the two sides are compared on.

The receipt was also unbound to what it scored: two arbitrary files were
accepted and reported under the default candidate and incumbent labels, and the
registry loader bypassed `TargetRegistry.from_json`'s format and content-hash
checks. Both artifacts are now verified against required digests before a byte
is read and recorded with those digests, the register loads through its
validating loader, and the per-target drift table the score block claimed to
offer is actually emitted. The seam driver's frozen-register check now compares
the same content hash through the same loader, so one artifact serves both
tools.

`uk_target_surface` claimed candidate and incumbent target surfaces agree
exactly. They are not independently sourced: the frozen parity instrument
carries input-column shares and no target surface, so the reference side is the
declared register and the candidate side the solve's realized diagnostics. That
is a real invariant — every declared target bound at its declared period — but
it is not the incumbent comparison the note promised, and the note now says so.
Gate manifest and spec-fingerprint pins recomputed for the edit.

The runbook pointed operators at the June builder, which constructs the
calibration stage with neither the measure resolver nor the exclusion register
and aborts on the first unmaterializable reference. It now documents the seam
driver, whose diagnostics digest is measured from the written bytes rather than
declared on the command line, and the scorer invocation that follows it.

Two findings reviewed and not reproduced, with the reasons recorded where a
reader would look: the person-anchor weight lookup cannot drop a duplicate
household id (the kernel validates group-table ids unique, and the CGT clone
offsets its clones' ids), and measure injection cannot mutate the source frame
(`UKFrameTargetAdapter` copies each table at construction) — the latter now
held by a test on the frame itself rather than on the restored output.

Reported by @MaxGhenis (findings 1, 2, 6, 7) and @vahid-ahmadi (1, 4).

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
The contract test declares its pins independently and a lockstep test holds
the two equal, so a re-cut moves both. Caught by running the data shard's UK
slice rather than only the build shard's pin suite.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
The generator validates every contract binding against its own closed
vocabulary, so declaring `value_reduction` and `band_filter_dimension` in the
contract without widening it there refuses regeneration outright.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
@juaristi22

Copy link
Copy Markdown
Collaborator Author

Thanks — dispositions on the second pass. Two confirmed and fixed, two not reproduced, and in both of those the shape that misled you is now commented where a reader looks.

1. _aggregate_admin_totals clone collapse — not reproduced

household_id is a group-table id column, and the Frame kernel refuses a bundle whose group ids are not unique before any frame exists (microcosm/frame/bundle.py:218-222). Probed directly:

ValueError: Group table 'household' id column 'household_id' must be unique;
duplicated ids include [0, 1].

The CGT clone is built for that rule rather than against it: cgt_structure.py:329-331 offsets every cloned household id by a multiplier and re-links its people, so clones are new households carrying half mass each, not repeated keys. dict(zip(..., strict=True)) therefore cannot drop a clone's weight.

You were right that strict=True does not check key uniqueness — it just isn't what guarantees it here. The guarantee is upstream and invisible at the call site, which is the part worth fixing, so the lookup now cites it.

2. _band_lower_edge not tied to the groupby variable — confirmed

Not live on today's data (each banded spec resolves exactly one band-like Ledger filter, so the alphabetical scan has nothing to get wrong), but the guess was real and its failure mode is exactly as you describe: a wrong-variable slice yields plausible, distinct, wrong subpopulation totals — strictly worse than the identical-across-bands defect 8eef7eb4 fixed, because that one was loud once measured.

The chronicle dimension names and the model's groupby_variable are not the same vocabulary (the fact's layout.groupby_dimension is a table-line id like hmrc.cgt_table1_line, not total_income), so binding by name alone would only work by coincidence. What it does now:

  • one band-like filter → use it, as before;
  • several → take the one whose key belongs to the binding's groupby_variable, or to a band_filter_dimension the binding declares when the publisher's and the model's spellings differ;
  • still ambiguous → raise, naming every candidate and the flag that resolves it.

Silence became either the right edge or a refusal. Three tests: right-edge-not-alphabetical-order, the refusal, and the declared tie-break.

3. UKMeasureResolver.knows unreachable — confirmed

Vacuously true for person/benunit/household, so the loop's provider does not know … refusal could never fire. It now answers for the routes compute_uk_measure_input actually takes: native; person↔group by broadcast or any-collapse; group↔group only for numerics, via map_to. A benunit-native categorical asked for at household grain gets the fence's message instead of a wrapped provider exception — and, as you say, that quality matters more now that the loop raises where the harness excluded with a receipt.

4. _inject_measure_inputs no-mutation claim — not reproduced

UKFrameTargetAdapter.__init__ deep-copies: self.tables = {entity: frame.table(entity).copy() for entity in frame.entities} (uk_runtime/ledger_targets.py:139), same for the link tables. So injection writes to per-round table copies and the docstring's claim holds — but you are right that nothing was holding it. test_seam_never_modifies_data_variables runs after restore() and would pass under aliasing too.

There is now a test that inspects the source frame's columns while the resolution loop probes, so the property is held where it lives rather than where it survives, and the docstring names the constructor copy as the thing it rests on.


Also in this push: your earlier F2/F4 remain as dispositioned, and the scorer they touched has been reworked further after @MaxGhenis's pass — it now materializes the register on both sides and authenticates both artifacts against required digests, which is what made the holdout question worth settling in the first place.

@juaristi22

juaristi22 commented Aug 25, 2026

Copy link
Copy Markdown
Collaborator Author

Thanks for the probes — having each claim runnable made this fast to adjudicate. All eight reproduced. Seven are fixed in this push; the eighth is a real defect in the June driver, which the runbook no longer sends anyone to and which #757 retires. Two further items are doctrine calls I've handed back to @juaristi22 rather than deciding mid-PR (will be addressed in the follow-up #757).

Branch is also merged up to post-#747 main, and the UK spec digest was recomputed rather than taken from either side: 0c85845b4d463638ae3e5c5a25e17de8b720794e3653c5991dc4f069d95762d3.

3 — boolean counts published as household indicators — confirmed, fixed

The one you flagged as most substantive, and it was. Eight contract rows declare a count of children over a boolean member variable; the crosstab read that boolean at the target's own grain, so the resolver's person→household route collapsed it with any and a household indicator went to the solver against a child-count target.

Counts now declare the reduction:

"value_variable": "is_child",
"value_reduction": {"variable": "is_child", "entity": "person", "reduce": "sum"}

UKFrameTargetAdapter grows entity_reduction, the numeric sibling of household_condition (same linkage, same membership guard), and an adapter that cannot reduce refuses rather than falling back to the same-grain read. The flag keeps its any semantics — for children_claimant_pip the same variable is both flag and value, which is precisely why the reduction has to be declared per role rather than inferred from dtype.

Your inventory of eight is exactly the set. Worth noting what made it survive review: the stub adapter and the fixture frame both carried a comment asserting the mapping summed ("uc_is_child_limit_affected sums to the number of flagged children, so it is both the affected flag and the affected-children count"). Both now carry real person rows where the flag (1/0/1) and the count (3/0/2) differ, so the distinction is exercised instead of asserted.

4 — unloggable seam — confirmed on all four sub-points, fixed

  • Scope: uk-national-calibrationuk/national, unratified. Renamed to uk-frs-calibrationuk/frs, joining the chain the spine and staging stages already share. Held by a test that derives the scope through tools/logbook.py and asserts it is declared.
  • Failure recording: the attempt is now wrapped; every terminal disposition writes an error receipt and a failed row. Your CLI/environment-predecessor scenario is covered by a test asserting a refusal leaves a row and no staged H5, diagnostics or gate report.
  • Predecessor ordering: resolve_predecessor moved ahead of every side effect, matching the existing UK driver.
  • Build id: was deterministic per release id, which both the local chain and the store reject on a re-run. Attempts carry unique ids now; the release id travels in run_config so the tie to the release survives.

1 — the documented command aborts — confirmed, fixed

The runbook now documents tools/calibrate_uk_national_dataset.py, and says why in one line: it is the only driver that builds the resolver and applies the committed exclusion register, and 187 of the activated references bind model outputs no frame carries.

2 — the rule-1 scorer — confirmed, mostly fixed

Reproduced. Both sides now go through the same resolve→inject→materialize route the seam takes (prepare_uk_target_frame), and a skipped target refuses rather than shrinking the surface the comparison runs on — there is a test with a production-shaped dwp/uc/households spec asserting exactly that refusal. Both artifacts are verified against required --candidate-sha256/--incumbent-sha256 before a byte is read and recorded with those digests; the register loads through TargetRegistry.from_json (format revision and content-hash re-derivation), and the seam driver's frozen-register check now compares the same content hash through the same loader, so one artifact serves both tools. The per-target drift table is emitted.

Not fixed here: cross-pinning the score receipt into signed run evidence. Doing that honestly means deciding what a combined spine-plus-calibration certification is, which is release-cut work (#757 work package B). The seam says shippable: false with that reason until then.

5 — operator-selectable release-candidate membership — confirmed, fixed in part

--release-candidate now refuses --measure-exclusions, for the reason you gave: the register prunes the compiled surface before the solve and the scoped battery carries no target-surface gate, so nothing downstream could see the narrowing. Same shape as the doctrine-override refusal @juaristi22 ruled on. The applied receipt no longer drops tracking, and the loader now requires it.

The full approver/adjudication/expiry schema you compared against weighted_integrity.py is a register-contract decision, not mine to make mid-PR — flagged below.

7 — the "independently sourced" parity trio — confirmed, corrected

Two comparisons travel in one evidence object and they are not equally strong. The column surfaces genuinely are independently sourced; the target surfaces are not — EfrsParityReference carries input-column shares and no target surface, so the reference side is the declared register and the candidate side the solve's realized diagnostics.

That is still a real invariant (every declared target bound at its declared period, at name@period grain — it catches a solve that silently drops rows), but it is not the incumbent agreement uk_target_surface claimed. The gate note now states what it proves and what it does not; the docstring says the same at the call site. Gate manifest and spec-fingerprint pins recomputed for the note edit: 4f66eea7… and 26fdbcccf…, method verified by reproducing the pre-edit pins from HEAD's gates.json first.

8 — verified-then-discarded Ledger identity — confirmed, fixed

run_uk_calibration accepted ledger_artifact and never used it, so digests the CLI had just verified reached neither the identity digest, the build record, nor the Logbook row. The verified identity (facts sha, manifest sha, row count, artifact id and profile) is now sealed into run_config; a bare feed's missing manifest is recorded as absent rather than invented.

6 — diagnostics digest signed before the diagnostics exist — confirmed, and it is the June driver

Reproduced, and it is worth separating: the new seam already does this correctly — it writes diagnostics, hashes the actual bytes, then constructs and signs the terminal evidence, and there is no --calibration-diagnostics-sha256 to supply. The defect is in tools/build_uk_national_dataset.py, which is the path #757 retires and which the runbook no longer sends anyone to. I have not changed it: it is the file this PR was scoped to leave alone, and the honest fix there is the same ordering rework the retirement does anyway. Flagging it rather than half-fixing a driver on its way out.


Two findings from @vahid-ahmadi's parallel pass did not reproduce (clone-collapse in the admin-anchor weight lookup; frame mutation through measure injection) — dispositioned in the sibling comment, with the reasons now written where each reader looked.

@juaristi22
juaristi22 merged commit bb796cb into main Aug 25, 2026
4 checks passed
MaxGhenis added a commit that referenced this pull request Aug 25, 2026
Main's #743 moved the UK bundle while this branch moves the shared seed
protocol; the merged UK digest combines both, BE carries only the
protocol move.

Refs #767

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
MaxGhenis added a commit that referenced this pull request Aug 26, 2026
…he union (third application)

Main moved the attested surfaces again (#743 first calibrated UK
candidate, #766 CI lane, #764 rename), so the merge re-pins the UK
spec_sha256, re-cuts the three gate-battery digests into the
microcosm-data contract and its test mirror, and regenerates the
release-input coverage manifest over the union - the d70ea39 pattern.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
juaristi22 added a commit that referenced this pull request Aug 26, 2026
…evidence

The calibration measure-exclusion register migrates to the weighted-integrity
record shape: approver, adjudication, canonical ISO approval and expiry dates,
with the window enforced when exclusions apply. Outside it the run refuses
with a correct-or-renew message naming the tracked gap, so a narrowing of the
target surface neither lapses silently nor lives forever — the gap the #743
audit named. The five salary-sacrifice entries carry Maria's three-month
window under that adjudication.

The owned_land input-mass exclusion was due to expire 2026-09-20 on E5-era
evidence. The stability instrument re-ran on the 25-stage candidate — the E5
method, adapted to strip post-wealth stage columns, drop the stacked SPI and
CGT rows, and clamp age to the stage-time top code — and the instability
persists: 53.8 percent national and 96.7 percent worst-region owned_land
swing between adjacent seeds, the uk-data#448 realization-variance class. The
exclusion re-signs on the fresh receipt with its one-month expiry and the
end-of-workstream revisit intact; the input-mass evidence pin and the
contract-test mirror follow the re-signed record.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

3 participants