UK national calibration run-readiness: doctrine, live parity trio, US-format diagnostics, scorer, staging posture (#623) - #743
Conversation
… re-map, ingest scale ladder The frozen release object uk/frs_release.json (survey/base 2024, calibration 2025, SN 9563, DOI, UKDS zip sha, HF acquisition pins) drives lockstep asserts over the re-pinned raw-tab manifest: all 21 frs_table artifacts re-pinned to the 2024-25 tabs with the SPI-convention keys, six stages' SN 9252 prose defect fixed, and the typed spec moved in lockstep. TIME_PERIOD and the HMRC SPI/CGT build periods move to "2024"; the 2023-24 HMRC published surface is re-mapped as a signed nearest-available-vintage declaration (period_mapping: latest_published_tax_year) with the frozen original byte-untouched and the source-contract validator reconstructing the live canonical payload. take_up_contract build_year 2024 flips the two 2024 date-keyed rates. The ingest driver joins the #627 scale ladder (--sample-fraction post-frs_spine via the generic frame-sampling helpers) with receipt-postures on three full-scale fences at sampled rungs, and the WAS bridge-donor locator defect that refused every full-roster licensed run is fixed with a hermetic pin-coherence regression test. Part of #723. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
The regenerator pins move to enhanced_frs_2024_25.h5 at the v1.56.14 tag (= incumbent-at-pin ebf733c) and the committed reference re-freezes at the new vintage: 145 columns (surface unchanged), period "2024", entity counts now test-pinned with the re-derived record-count identity (16,288 raw + 10,000 SPI) x 2 + 270 CGT band donors = 52,846 - no raw household is dropped at 2024-25 and the donor count follows the HMRC band file (30 x 9). The release-input coverage manifest regenerates against the new reference with the certified 2023 candidate unchanged (candidate side moves at #686); known gaps stay empty and the restored-column receipts hold. Part of #723. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
The efrs-post-calibration input-mass descriptor is replaced in place (same name, #687's replacement model): enhanced_frs_2024_25.h5 identity and the totals_sha256 of the regenerated 131-column weighted-totals evidence, with the registry, gates.json, and the data-shard publication mirrors re-pinned in the same reviewed change per the gate-battery contract. No thresholds move - #723 records the re-measured baselines, #686 arms them. UK_REFERENCE_DATASET_NAME follows the incumbent's 2024-25 dataset name. Both per-reference reviewed exclusions are re-signed against the new reference (charitable_investment_gifts; owned_land on a fresh 2024-25 stability receipt), pending the approver's confirmation on the PR. Part of #723. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
…ise channel-blind scope note UK_REFERENCE_DATASET_NAME drops the _recalibrated suffix (adjudicated 2026-08-20): the pinned reference is the published enhanced_frs_2024_25.h5 itself and no recalibrated variant exists at this vintage, so the gate-report label now names the artifact exactly (June report strings keep their own label). The registry scope notes stop claiming the incumbent "structurally lacks" the SPI clone channel - the 2024-25 artifact carries the synthetic rows structurally but no admin-restored mass in the channel-exclusive columns, which is the fact the reviewed exclusions rest on; the approved exclusion reasons are untouched. Gate digests re-cut over the post-#729 union in the same change. Part of #723. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
…block; self-describing June freeze Finding 1 (accepted): the family-coverage block rendered the #723 re-mapped period fields from the canonical manifest while hashing only the frozen mirror - evidence fields and their hash must name the same bytes. The block now carries a dual pin (source_manifest for the frozen June identity, canonical_source_manifest for the bytes the re-mapped fields come from) and a test binds each field set to the sha256 of the file it actually derives from. Finding 2 (rejected with armor): the committed replay report is the June evidence freeze and deliberately keeps mapped_build_period 2023 - it is evidence for the grandfathered release, not the 2024 line, and retires with the frozen manifest after #686 per #687; instead of regenerating it, a new assertion binds it to the FROZEN manifest's declared mapping so the partition is self-describing. Part of #723. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
…ibrated-candidate-623
…tree Both parents edit uk/gates.json and uk/country_package.json, so the merged manifest digests differ from either side's pins. Re-pinned by recomputation (never by picking a side): spec bundle e12a2cb8…, policy 404968fb…, gates manifest 59c7808d…, spec fingerprint bfb98736…. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
…-format diagnostics, scorer, staging posture Work items 1-7 of the narrowed #623 plan (runs held per the 2026-08-21 adjudications; everything here is synthetic/hermetic): - uk_runtime/national_doctrine.py: every solver constant the national calibration inherits becomes a declared, tamper-tested doctrine value (epochs 256, lr 0.02, ratio 10.0, seed 0, loss cap 10.0, l0 0.0, free mass, uniform weights). - Single compile path: the stage consumes the driver's compiled registry at the release calibration year and materializes bindings through the shared target_materialization interpreter; prepared scratch columns are restored post-solve, so the staged frame survives the real HDFStore writer (closes #729 dispositions finding 4). - Canonical mass/weight conventions mirroring the rowwise path: declared mass reason, post-solve fence (CALIBRATED kind, exactly one appended record), household_weight_kind_chain + calibration_mass_change manifest. - The parity trio evaluates for real on armed builds: candidate side from the staged frame and solve diagnostics, reference side from the frozen eFRS parity instrument and the declared registry at name@period grain — never a copied reference; unarmed builds keep evidence_absent. - calibration_diagnostics.json is now produced by the same shared diagnostics producer the US release path uses (per-epoch loss trajectory, per-target rows), with a pinned US<->UK format-parity test; the build block carries the chronicle artifact provenance (facts/manifest shas, profile ids) and reserves score_vs_enhanced_frs. - tools/score_uk_national_candidate.py: #578 rule-1 scoring, both artifacts rescored on one frozen register, June-schema score block with a declared none_declared holdout basis. - Declared non-certified staging-candidate input posture (sha-gated with a mid-read race guard, refused for release candidates) and the held-run runbook, unblocking on WS-E E8+E10 or an explicitly named base. Implemented by Codex from the reviewed plan; review pass fixed the materialization period (declared calibration year, never the frame's base-year time_period) and the parity-evidence sourcing above. Part of #623 under #665. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
…d adapter Both PRs merged upstream, including María's adaptations of parts of this branch's work (c534517 routes calibration through the shared materializer; 1d066e5 is the cherry-picked #733 review fix). Reconciliation: - target_materialization.py and verify_uk_identity_stability.py: main's reviewed versions taken wholesale. - national_calibration.py: this branch's registry+period+doctrine stage kept (all tests and the driver target it); its private frame adapter replaced by a lifecycle subclass of main's shared UKFrameTargetAdapter — one adapter, now writer-safe (prepared scratch columns restored away post-solve, per the adjudicated materializer-owned lifecycle). - Main's packaged-binding stage tests ported to the registry+period API; the persistence assertion inverted to the adjudicated writer-clean invariant (result columns exactly equal the input columns); the materialization stub adapter adopts the shared count-variable convention. - Digest pins auto-merged to main's post-union re-cut and verified green by the pin suites — no re-measurement needed. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
…ain merge The post-#735/#733 merge at 9a8bd46 resurrected the standalone `except ValueError` block that #733's review commit 1d066e5 had folded into the single generic handler. Because the narrower clause catches first, the named-edge branch inside `except Exception` became unreachable, silently reverting a reviewed structural fix. This branch has no business touching the spine driver at all — its scope is the national calibration seam — so the file returns to main's version exactly. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
An armed UK national build crashed before its first stage: _source_pins stored the full Ledger provenance block under the 'ledger_facts' role, but role_pins_digest requires each role to carry exactly sha256 and size_bytes, so every armed run raised ValueError at input_pins_digest. The two halves landed independently — the strict pin contract with Logbook adoption (#666) and the Ledger role pin with the compile-parity wiring — and neither PR branch fires it alone, because only an armed run (--ledger-facts) reaches this path and the licensed runs were held. It reproduces on both #743 and #747 branches and on main. The pin now carries the feed's verified digest and byte size; the richer Ledger identity block already travels in safe_artifacts, source_vintages, and the diagnostics build block, so nothing is lost. This fix belongs upstream in #743 (the run-readiness PR whose runbook documents the armed command); it rides the build branch until then.
… production path First armed run receipt: 187 of 388 references skipped at materialization, because they bind simulated tax-benefit outputs (income tax, NICs, UC, caseloads, payment-band crosstabs) and the calibration stage's frame adapter reads stored columns only. UKPolicyEngineAdapter — the live-sim adapter — has no production caller, and the stage's tests feed synthetic frames with measure columns pre-attached, so every real input (spine or certified June candidate) fails the 388-target surface the same way. The June release never hit this because its 149 targets were demographics-only. The runner now drives the materializer's own skip reports: each round computes the missing (entity, variable) pairs from a live policyengine-uk simulation over the same records — native-entity calculate, numeric map_to, categorical group-to-member broadcast, boolean any-collapse — attaches them to the frame, and retries until the register binds. The incumbent gets the same treatment before scoring, so rule 1 compares both datasets under one yardstick, and the score block declares that symmetry. References this posture cannot bind (the salary-sacrifice counterfactual deltas need adapter.counterfactual_delta) are excluded with a per-reference receipt. The production fix belongs in the #743 lane: either wire UKPolicyEngineAdapter into the stage or land this materialization as a declared pre-calibration step.
An armed UK national build crashed before its first stage: _source_pins stored the full Ledger provenance block under the 'ledger_facts' role, but role_pins_digest requires each role to carry exactly sha256 and size_bytes, so every armed run raised ValueError at input_pins_digest. The two halves landed independently — the strict pin contract with Logbook adoption (#666) and the Ledger role pin with the compile-parity wiring — and neither PR branch fires it alone, because only an armed run (--ledger-facts) reaches this path and the licensed runs were held. It reproduces on both #743 and #747 branches and on main. The pin now carries the feed's verified digest and byte size; the richer Ledger identity block already travels in safe_artifacts, source_vintages, and the diagnostics build block, so nothing is lost. This fix belongs upstream in #743 (the run-readiness PR whose runbook documents the armed command); it rides the build branch until then.
243 of the 388 UK references are banded, and none of them were ever sliced. Every employment-income band materialized 35,351,186 (roughly everyone with employment income) and every state-pension band 13,518,447 (roughly every state pensioner), so a band's apparent overshoot -- up to 6,759x on the state-pension 1m+ cell -- was an unsliced total compared against a band value, not a data defect. The slice was declared but unconsumed: the contract binding names groupby_variable, each compiled spec carries its own band's lower edge in Ledger filter metadata, and _prepared_column_values read neither. The binding's own filters list did work (verified: a family_type == SINGLE binding correctly masks a COUPLE row), so this is specifically the band dimension. Both published encodings reduce to one lower edge -- a numeric *_lower_bound (HMRC SPI) or a range label in monthly units scaled by band_period_factor (DWP awards) -- because no reference anywhere declares an upper bound. A band's upper edge is its sibling's lower edge within the same contract target, grouped per contract target rather than per dimension so two measures sharing a dimension cannot slice each other on the wrong boundaries; the top band runs to infinity. Validated against the real register: 243 banded references, zero unreadable, and the derived bounds match the measure names (150000-200000, 1000000-inf, and UC monthly 500.01-600.00 to annual 6000.12-7200.12). A band whose edge cannot be read now raises, so the measure is skipped and reported rather than silently reporting the whole population as one band -- the failure mode that hid this. Tests cover the partition property, an adjacent-bands-differ regression, the monthly label conversion, and the refusal. Lives in the shared materialization module, so it cherry-picks into any branch.
All 15 two-child-limit references failed to materialize. Two naming faults,
both in the contract rather than the provider:
- eight children references declared value_variable "children_count", which
is not a policyengine-uk variable. They now bind what they actually count:
uc_is_child_limit_affected for the affected-children rows (mapped to
household it sums to the count of flagged children) and is_child for the
children-in-affected-households rows (the total children there).
- fifteen bindings put a prose label in count_of ("affected_households",
"affected_children", "children_in_affected_households"). count_of is a
column fallback consulted when value_variable is an entity-count
indicator, so the provider looked those labels up as columns and raised.
The labels move to notes, leaving household references on the
household_count unit indicator as intended.
The provider keeps its documented behaviour; the pre-existing test covering
count_of as a real column still passes.
The stub adapter and national-stage frame both modelled the old children_count column. Mapped to household, uc_is_child_limit_affected sums to the number of flagged children, so it serves as both the affected flag and the affected-children count; the fixtures now carry counts rather than indicators. Caught by running the suites properly: the earlier chain piped pytest into tail, so its exit status was tail's and two real failures rode through into 760fe8b.
UKFrameTargetAdapter.household_condition built a group-membership column for every non-household entity, so a condition declaring entity "person" asked for "person_person_id" and raised. People sit directly in a household; only group entities need that lookup. Every ONS household-composition reference uses person-level conditions (is_child, age), so all ten were excluded from calibration. That left household structure unconstrained while population stayed targeted, and the solve satisfied population by inflating households rather than multiplying them: weighted benunits per household reached 1.476 against the incumbent's 1.132, from an identical unweighted 1.158. The downstream cost is Universal Credit. 75.5% of the candidate's single UC claimants end up in multi-benunit households (incumbent: 8.6%), where only one benunit claims housing costs, so the housing element reaches 47.2% of them against the incumbent's 73.9% -- despite the candidate having more renters (88.3% vs 79.2%) and higher rent, and despite beating the incumbent within both strata (92.6% vs 78.9% solo, 32.5% vs 20.4% multi). Simpson's paradox, driven entirely by composition. Median single UC award falls to GBP 8,059 against GBP 9,310, emptying the GBP 8-11k and GBP 14-19k award bands that carry 98% of the caseload shortfall. Shared UK runtime, so this cherry-picks into the spine lane.
The 1a3274b crash fix landed without a test; this locks the contract it restored: the ledger_facts role pin carries exactly {sha256, size_bytes} (both feed layouts) and survives role_pins_digest, while the full provenance block is rejected — the exact shape that crashed the first armed run at input_pins_digest. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
…ine rule into the solve Re-cut of the build branch's 035ed63 under María's 2026-08-24 ruling: family_equal enters _ALLOWED_TARGET_WEIGHT_RULES with its armed-run receipts, and uk_national_target_loss_weights() derives the family-share vector — but the doctrine DEFAULT stays uniform. She passes family_equal as an explicit per-run setup while the weighting doctrine is measured; neither rule is adopted by default (run-9 receipts cut both ways). Also fixes the latent wiring defect the armed run exposed: the stage echoed target_weight_rule in its manifest but never passed a weight vector to calibrate(), so any declared rule silently solved uniform. The vector now travels explicitly; uniform maps to None, keeping the shipped identity byte-stable under the default. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
The two-child-limit rebinding (760fe8b) replaced the prose count_of labels with real value_variable columns; the chronicle loader-guarantee test still demanded count_of on every baseline_flag_crosstab binding. The guarantee now matches the provider: affected_flag_variable plus either counted-column spelling. Latent on the build branch too — this suite was never re-run there after the rebinding. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
|
Automated review pass (Claude Code, high effort, diff only — no verification runs). Posting on the draft since two of these bear on whether the calibration gate means anything; the other two are dead code. Substantive1. 2. Dead code3. 4. Findings 3 and 4 are trivial. 1 and 2 are the ones I would want resolved before this leaves draft — in both cases a gate reports success without having tested the thing it names. |
…d measure-resolution loop Two moves the calibration seam needs, both behaviour-neutral for the June path. The H5 reader/writer pair moves next to uk_national_frame and validate_uk_national_frame — its natural home — with national_build re-exporting every name, so the legacy driver and its suites are unchanged. The seam can then load and stage frames without importing the legacy build module at all. target_materialization grows the country-agnostic half of the assessment harness's resolution loop: probe-adapter rounds, once-per-round dedup of a shared missing key, and a fail-loud finish. The country-specific step is injected as a provider, so nothing about policyengine-uk enters the shared module. Where the harness excluded an unresolvable reference with a receipt, production raises: register pruning belongs to a reviewed exclusion register, never to a silent loop. assert_calibration_input_finite is the seam's NaN fence — it lists every offending column at once instead of failing on the first. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
…overrides 187 of the 388 compiled references bind model outputs — income tax, UC — that no production code ever computed onto a frame; the stage's adapter reads stored columns only, which is why the first armed run could not bind them. The provider computes them through policyengine-uk on four receipted routes and injects them at adapter level, where names are table-scoped, so the Frame's global column-uniqueness rule is never touched and the staged artifact carries no scratch column. The five salary-sacrifice counterfactual references cannot bind without a live counterfactual run, so they leave the register through a reviewed exclusion file with a written reason each, not through a silent skip. A stale entry that matches nothing is an error. Doctrine defaults stay the reviewed constants. uk_doctrine_with_overrides builds an effective doctrine through the frozen dataclass, so the closed vocabularies still apply, and returns the diff against v1 — recorded as an explicit deviation wherever the run is described. Only the four fields calibration work actually turns are overridable; seed, weight ratio, mass and scale rules stay reviewed constants and refuse. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
…re given The June national driver rederives SPI income, redraws capital gains and replays support tables onto its input — work the WS-E spine now ships already derived, and which its own zero-weight precondition refuses to perform on a spine at all. This seam does none of it: it verifies the input pin, loads, fences NaNs, solves weights under doctrine, and writes the staged H5, US-format diagnostics, build record and signed gate report. Its defining invariant has a test: every data column of the staged artifact is identical to the input, and only the household weights, their kind and one appended mass record differ. Nothing here imports the legacy build module, so the June path can retire on its own schedule without this lane moving. The gate posture is calibration-scoped. Spine-construction gates cannot pass a build that constructs no spine, and pretending otherwise would be the invented pass this repo refuses; each out-of-scope entry is listed in the signed report with its reason, and a test asserts scope and exclusions together cover every declared gate, so a new gate cannot land unclassified. Aggregate-admin evidence follows the anchors' own convention — per-household means for the NEED anchors, a national total for NHS — and refuses when the frame cannot supply an anchor rather than reporting a zero. The build record states shippable: false with its reason. A publishable artifact needs the full battery, which is release-cut work. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
…ntage line Finding 2 (confirmed). The score block copied final_loss into candidate_holdout_loss and incumbent_holdout_loss while declaring holdout_basis none_declared. The June fixture this block mirrors carries a genuinely different holdout value — 0.1239 against a 0.0159 train loss, the eightfold degradation that made its holdout the discriminating comparison — so repeating the fitted loss there reads downstream as perfect generalization from a measurement never made. The keys now report absence, and declaring a basis without computing its split raises rather than falling back. The scorer also now takes the H5 loader from its new home instead of the legacy build module. Finding 4 (confirmed). source_vintages["frs"] was assigned twice, so load_uk_frs_release() ran twice; main carries one line, which makes this a branch-local artifact of the same merge that duplicated the spine driver's failure handler. A sweep of tools/ and uk_runtime/ for the same defect class found no others. Finding 1 (refuted, comment only). The staging-tier early return skips the certified-pin equality, which is what the tier declares, but the bytes are bound before it: the verification token is a module-private sentinel only the two verifiers stamp, and the staging verifier hashes the file against a mandatory declared sha with a mid-read race guard. The comment now says so, since the shape invites the reading twice now. Finding 3 was already fixed on this branch before the review was posted. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
|
Thanks — dispositions below. Two confirmed and fixed, one refuted with evidence, one that was already fixed on the branch before the review landed. A timing note first: the review read origin's then-head 2 — scorer holdout comparison: CONFIRMED, fixedYou're right that the block reports a holdout it never measured, and the June fixture makes it concrete.
June held rows out, and its holdout was the discriminating comparison — an eightfold train-to-holdout degradation on the candidate. Copying Fixed by reporting absence instead of a fabricated number: both holdout keys are now Two refinements to the finding. "Can only ever be won by the candidate" is stronger than what holds — the scorer's own synthetic-twins test has the candidate losing on full loss (0.25 against the incumbent's 0.1), and the candidate's structural advantage is declared, not hidden, in 4 — duplicated
|
Scope update: the calibration seam now lands here, with the run campaign's receiptsThis PR opened as run-readiness tooling with the licensed run held. Since then the first armed campaign actually ran (9 runs on the Campaign fixes (upstreamed from the build branch)
Deliberately not picked: the epochs-1500 default (same ruling), and The calibration seam (new scope)The production driver rederives SPI income, redraws capital gains and replays support tables onto its input — work the WS-E spine now ships already derived, and which its own zero-weight precondition refuses to perform on a spine at all. So a spine-based calibration was not a matter of flags; there was no code path. Per María's architecture ruling the seam does not wrap
VerificationFull per-shard suite green on the pushed tip — frame, fit, calibrate, build and data all exit 0 — plus Explicitly not hereRegistry coverage: 29 contract targets carry no reference row and so never compile. That is verified and written up as an adjudication package in #736 — it needs several rulings and a digest re-cut, so it does not ride this PR. The first production calibration run is likewise a separate step, once this and #747 land. 🤖 Generated with Claude Code |
…ibrated-candidate-623 # Conflicts: # packages/microcosm-build/src/microcosm/build/uk/country_package.json # packages/microcosm-build/tests/test_country_spec.py # packages/microcosm-build/tests/test_spec_engine_country_bundles.py
The wheels gate installs the base wheels without the optional pandas HDF backend, so the three calibration-run tests that write and re-read a real staged H5 failed there with an ImportError while passing everywhere else. They now carry the repo's existing guard, matching every other test that touches the HDF backend. The two other new seam test files monkeypatch the writer rather than performing real H5 I/O, so they need no guard and already passed the gate. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
|
Second pass (Claude Code, high effort, diff only — no build or test execution), scoped to the new material: the calibration seam ( First, on the earlier round: the refutation of finding 1 is accepted and it was my error. The reviewer read the early return without tracing the guards above it — Four findings on the new scope. 1.
|
|
Ran a deep automated review of this PR through my agent session before your merge (sol lane, adversarial-probe style; CI green is acknowledged — these are checks below the CI waterline). Verdict from the run: hold for five findings, each reproduced by a committed probe rather than asserted. Full code-cited report: REVIEW-743.md on The five, ranked:
Your call entirely — if any of these misread the code, say so and merge; the probes and their outputs are in the report so each claim is checkable in minutes. Happy to pair on fixes or have my agents draft them if useful. (#3 seems the most substantive: it would bias the UC family-composition targets the #735 surface just activated.) |
…ibrated-candidate-623 # Conflicts: # packages/microcosm-build/tests/test_spec_engine_country_bundles.py
…band Three review findings on the calibration seam's measure surface, each one a wrong number that no error would have surfaced. The two-child-limit contract declares eight rows whose value is a count of children over a boolean member variable. The crosstab read that boolean at the target's own grain, so the resolver's person-to-household route collapsed it with `any` and published a household indicator against a child-count target. Counts now declare `value_reduction`, the adapter grows the numeric sibling of `household_condition` to sum members, and an adapter that cannot reduce refuses instead of falling back to the same-grain read. The stub and fixture frames that encoded the old assumption in a comment now carry real person rows, so the distinction is exercised rather than asserted. Band edges were taken from the alphabetically-first band-like Ledger filter with no tie to the binding's groupby variable, so a spec carrying two banded dimensions could slice on the wrong one — distinct, plausible, wrong numbers, the same class as the unsliced-measure defect one layer up. The edge is now attributed to the declared banding dimension, `band_filter_dimension` names it where the publisher's and the model's spellings differ, and an unresolvable tie raises. `UKMeasureResolver.knows` was vacuously true for every entity, so the resolution loop's refusal could never fire and an unmappable categorical surfaced as a wrapped provider exception. It now answers for the routes `compute_uk_measure_input` actually takes. Reported by @vahid-ahmadi (2, 3) and @MaxGhenis (1). Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
…hat it verified The seam declared pipeline `uk-national-calibration`, which derives the scope `uk/national` — not in the ratified vocabulary. The FRS line's spine, staging, imputation and calibration stages share one chain by design (the dataset token names the base data, not the build mechanism), so calibration is `uk-frs-calibration` and lands on `uk/frs`. Three more lifecycle gaps went with it. The predecessor digest was resolved at the very end, so a disagreeing chain refused only after a staged H5, diagnostics and a signed gate report already existed; it is now validated before anything is written. Only successful runs recorded a row, against the binding rule that successful, failed, refused and discarded attempts all produce one; the attempt is now wrapped, and a refusal writes an error receipt and a `failed` row while staging nothing. And the build id was deterministic per release id, which both the local chain and the store reject on a re-run; attempts now carry unique ids and the release id travels in run_config. `run_uk_calibration` also accepted `ledger_artifact` and never used it, so the facts and manifest digests the caller had just verified reached neither the identity digest, the build record, nor the Logbook row. The verified identity is now sealed into run_config; a bare feed's missing manifest is recorded as absent rather than invented. Release candidates no longer accept `--measure-exclusions`: the register prunes the target surface before the solve and the calibration-scoped battery carries no target-surface gate, so an operator file could narrow what a candidate was measured against unnoticed. The applied receipt now keeps each exclusion's tracking reference, and the loader requires one. Reported by @MaxGhenis (findings 4, 5, 8). Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
…oves The scorer loaded both H5s and called `score_targets` directly. Every packaged UK reference binds a slash-named prepared measure, and calibration deliberately strips those before export — so on a production register it bound nothing and aborted, while its tests stayed green on plain `measure_a` columns no register uses. Both sides now go through the same resolve-inject-materialize route the calibration seam takes, via `prepare_uk_target_frame`, and a skipped target refuses rather than quietly shrinking the surface the two sides are compared on. The receipt was also unbound to what it scored: two arbitrary files were accepted and reported under the default candidate and incumbent labels, and the registry loader bypassed `TargetRegistry.from_json`'s format and content-hash checks. Both artifacts are now verified against required digests before a byte is read and recorded with those digests, the register loads through its validating loader, and the per-target drift table the score block claimed to offer is actually emitted. The seam driver's frozen-register check now compares the same content hash through the same loader, so one artifact serves both tools. `uk_target_surface` claimed candidate and incumbent target surfaces agree exactly. They are not independently sourced: the frozen parity instrument carries input-column shares and no target surface, so the reference side is the declared register and the candidate side the solve's realized diagnostics. That is a real invariant — every declared target bound at its declared period — but it is not the incumbent comparison the note promised, and the note now says so. Gate manifest and spec-fingerprint pins recomputed for the edit. The runbook pointed operators at the June builder, which constructs the calibration stage with neither the measure resolver nor the exclusion register and aborts on the first unmaterializable reference. It now documents the seam driver, whose diagnostics digest is measured from the written bytes rather than declared on the command line, and the scorer invocation that follows it. Two findings reviewed and not reproduced, with the reasons recorded where a reader would look: the person-anchor weight lookup cannot drop a duplicate household id (the kernel validates group-table ids unique, and the CGT clone offsets its clones' ids), and measure injection cannot mutate the source frame (`UKFrameTargetAdapter` copies each table at construction) — the latter now held by a test on the frame itself rather than on the restored output. Reported by @MaxGhenis (findings 1, 2, 6, 7) and @vahid-ahmadi (1, 4). Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
The contract test declares its pins independently and a lockstep test holds the two equal, so a re-cut moves both. Caught by running the data shard's UK slice rather than only the build shard's pin suite. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
The generator validates every contract binding against its own closed vocabulary, so declaring `value_reduction` and `band_filter_dimension` in the contract without widening it there refuses regeneration outright. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
|
Thanks — dispositions on the second pass. Two confirmed and fixed, two not reproduced, and in both of those the shape that misled you is now commented where a reader looks. 1.
|
|
Thanks for the probes — having each claim runnable made this fast to adjudicate. All eight reproduced. Seven are fixed in this push; the eighth is a real defect in the June driver, which the runbook no longer sends anyone to and which #757 retires. Two further items are doctrine calls I've handed back to @juaristi22 rather than deciding mid-PR (will be addressed in the follow-up #757). Branch is also merged up to post-#747 3 — boolean counts published as household indicators — confirmed, fixedThe one you flagged as most substantive, and it was. Eight contract rows declare a count of children over a boolean member variable; the crosstab read that boolean at the target's own grain, so the resolver's person→household route collapsed it with Counts now declare the reduction: "value_variable": "is_child",
"value_reduction": {"variable": "is_child", "entity": "person", "reduce": "sum"}
Your inventory of eight is exactly the set. Worth noting what made it survive review: the stub adapter and the fixture frame both carried a comment asserting the mapping summed (" 4 — unloggable seam — confirmed on all four sub-points, fixed
1 — the documented command aborts — confirmed, fixedThe runbook now documents 2 — the rule-1 scorer — confirmed, mostly fixedReproduced. Both sides now go through the same resolve→inject→materialize route the seam takes ( Not fixed here: cross-pinning the score receipt into signed run evidence. Doing that honestly means deciding what a combined spine-plus-calibration certification is, which is release-cut work (#757 work package B). The seam says 5 — operator-selectable release-candidate membership — confirmed, fixed in part
The full approver/adjudication/expiry schema you compared against 7 — the "independently sourced" parity trio — confirmed, correctedTwo comparisons travel in one evidence object and they are not equally strong. The column surfaces genuinely are independently sourced; the target surfaces are not — That is still a real invariant (every declared target bound at its declared period, at 8 — verified-then-discarded Ledger identity — confirmed, fixed
6 — diagnostics digest signed before the diagnostics exist — confirmed, and it is the June driverReproduced, and it is worth separating: the new seam already does this correctly — it writes diagnostics, hashes the actual bytes, then constructs and signs the terminal evidence, and there is no Two findings from @vahid-ahmadi's parallel pass did not reproduce (clone-collapse in the admin-anchor weight lookup; frame mutation through measure injection) — dispositioned in the sibling comment, with the reasons now written where each reader looked. |
…he union (third application) Main moved the attested surfaces again (#743 first calibrated UK candidate, #766 CI lane, #764 rename), so the merge re-pins the UK spec_sha256, re-cuts the three gate-battery digests into the microcosm-data contract and its test mirror, and regenerates the release-input coverage manifest over the union - the d70ea39 pattern. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
…evidence The calibration measure-exclusion register migrates to the weighted-integrity record shape: approver, adjudication, canonical ISO approval and expiry dates, with the window enforced when exclusions apply. Outside it the run refuses with a correct-or-renew message naming the tracked gap, so a narrowing of the target surface neither lapses silently nor lives forever — the gap the #743 audit named. The five salary-sacrifice entries carry Maria's three-month window under that adjudication. The owned_land input-mass exclusion was due to expire 2026-09-20 on E5-era evidence. The stability instrument re-ran on the 25-stage candidate — the E5 method, adapted to strip post-wealth stage columns, drop the stacked SPI and CGT rows, and clamp age to the stage-time top code — and the instability persists: 53.8 percent national and 96.7 percent worst-region owned_land swing between adjacent seeds, the uk-data#448 realization-variance class. The exclusion re-signs on the fresh receipt with its one-month expiry and the end-of-workstream revisit intact; the input-mass evidence pin and the contract-test mirror follow the re-signed record. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Part of #623 under #665 (WS-D). Registry #736 item 10 asked to close-or-narrow #623 after #729 delivered the seam core; this PR is the narrowing, per the adjudicated plan (María, 2026-08-21): all calibration tooling lands now — solve doctrine, honest gate evidence, US-format diagnostics, candidate-vs-incumbent scoring, run posture — and the licensed armed run itself stays held until the WS-E spine completes E8 (#684) + E10 (#686), or María names a different base. Builds on merged #735 (the 2025 Ledger surface) and #733 (the FRS 2024-25 line).
What lands
uk_runtime/national_doctrine.pymirroringlocal_doctrine.py: every solver option the seam previously inherited silently (epochs 256, lr 0.02, max_weight_ratio 10.0, seed 0, loss cap 10.0, l0 0.0, free mass, uniform target weights) is a declared, tamper-tested constant. Doctrine v1 deliberately lifts Add the UK national calibration step over ledger-backed target references #729's values; first-run evidence is the revision path (adjudication 3).UKNationalCalibrationStageconsumes the driver's compiled registry atload_uk_frs_release().calibration_year(2025) and materializes bindings through the shared interpreter via a lifecycle subclass of the sharedUKFrameTargetAdapter: prepared scratch columns are restored away post-solve, so the staged frame survives the real HDFStore writer (closes Add the UK national calibration step over ledger-backed target references #729 dispositions finding 4, per adjudication 2 — materializer-owned lifecycle). The materialization period is passed explicitly — never the input frame's base-yeartime_period, which lags the calibration year.national_calibration_mass_reason()names the bound families; a post-solve fence requires the CALIBRATED kind and exactly one appended mass record; the stage manifest carries the rowwise-conventionweightsblock (household_weight_kind_chain,calibration_mass_change, solve block, doctrine echo)._stage_parity_evidencebuilds theuk_export_surface/uk_target_surface/uk_target_fitevidence from independently sourced sides: candidate from the staged frame and solve diagnostics, reference from the frozen eFRS parity instrument (parity_referenceparameter, driver passes the committed instrument) and the declared registry atname@periodgrain. A copied reference can never fabricate a pass; unarmed builds keep the honestevidence_absent, pinned in both postures.uk_calibration_diagnostics_payload/write_uk_calibration_diagnostics— the same shared producer the US release path imports (schema v6, per-epochloss_trajectory, per-target rows, loss attribution), so the calibration dashboard consumes UK and US files identically. A hermetic format-parity pin asserts the UK payload's shared layer equals the shared producer's output withuk_diagnosticsstrictly additive. Thebuildblock carries the chronicle artifact provenance (facts sha256, manifest sha256, profile ids), build id, code pins, and input posture — every target value traceable throughledger_facts_sha256(Publish Chronicle package IDs in calibration diagnostics targets #661).tools/score_uk_national_candidate.py: both artifacts rescored on the same frozen register withrelative_error_loss(cap 10.0), emitting the June-schemascore_vs_enhanced_frsblock ({candidate,incumbent}×{train,holdout,full}_loss + target_wins) with per-family wins. Signed differences:holdout_basis: "none_declared"(June's holdout split lives only in the archived pipeline); the diagnosticsbuildblock reservesscore_vs_enhanced_frs: nulluntil the licensed run merges the real receipt.--staging-candidate-input-sha256: sha-gated with a mid-read race guard, labeled tier in the build record, refused for release candidates) so the armed run is push-button on any pre-clone spine;docs/uk-national-calibration-runbook-623.mdrecords the exact command, the evidence-dir layout, and the two unblock conditions. Grain basis (adjudication 1): the incumbent's publishedenhanced_frs_2024_25.h5is itself pre-clone (52,846 households;clone_and_assignfeeds only its local-weights product), so pre-clone scoring is apples-to-apples — with one signed method difference to carry into the score receipt: incumbent national weights are a collapsed local solve, ours is a direct national solve under doctrine.Explicitly out of scope
Verification
Hermetic: doctrine pinned tests; single-compile-path + writer regression through the real
write_uk_national_frame; armed/unarmed parity-trio postures; mass-record fences; US↔UK diagnostics format-parity pin; scorer unit tests on synthetic twins; staging-posture flag matrix (adversarial refusals). Synthetic end-to-end in CI: armed synthetic build → real writer → full battery, US-format diagnostics, Logbook row withledger_factsbound ininput_pins_digest. Full three-shard suite + ruff green locally on the merged tip.🤖 Generated with Claude Code