Deferred from #355 via /brainstorm 341 355 365 (decision-log entry 85); #355 is now closed — its policy half shipped as approved, CI-enforced spec 023-authoring-assurance (PRs #370/#374). This issue is the standalone owner of the empirical half. Activation trigger revised in decision-log entry 89 (/brainstorm blocked tickets, 2026-09-08).
Why deferred
Nothing is blocked on the evaluation. The policy (risk-tier stance, authoring.method provenance marker, split reviewer qualification — all in spec 023) gives reviewers a boundary today. The corpus + controlled run + measurement report is large work whose shape depends on there being enough real LLM-assisted publishes to measure.
Activation condition
Start this work when either:
- ≥ 5 published capabilities carry
authoring.method: "llm-assisted" — a one-line count over capabilities/**/contract.json on main. Progress: ~3 / 5 as of 2026-09-08 (audio.calibration-plan-create@1.0.0, inference.evidence-normalize@1.0.0, inference.evidence-normalize@1.0.1). #369 was the first.
- or there is an owner decision to move
023's low-risk deterministic tier from experimental to supported (this evaluation's data is the precondition for that move).
Activation means "begin building the corpus and methodology" — not "there is already enough data for a credible report". The report's own credibility bar (sample size, per-bucket n) is set by the methodology when it is written, the same split #365 uses.
Scope when activated (absorbs every unfinished #355 DoD item)
- A versioned evaluation corpus covering deterministic/pure, validation, effectful, and explicitly-excluded high-risk capability classes.
- Every generated candidate linked to its contract, source revision, artifact digest, test evidence, and human review decision (spec
023 FR-003 provides most of this chain).
- Report measures: first-pass validation rate, security/policy rejection rate, reviewer rework, defect escape, time-to-accepted artifact. Methodology and limitations published.
- Contract-derived tests + at least one property-based or metamorphic check where the contract admits them.
- The report's data sets the numeric bar for the
experimental → supported promotion of the low-risk tier (a later standalone owner decision, per 023 FR-004 / Governing Relationship).
(#355 DoD items "policy by risk tier", "non-Rust vs qualified review guidance", and "no bypass of existing gates" are already delivered in spec 023 FR-004 / FR-005 / FR-006 and are not repeated here.)
Definition of done
Related
Deferred from #355 via
/brainstorm 341 355 365(decision-log entry 85); #355 is now closed — its policy half shipped as approved, CI-enforced spec023-authoring-assurance(PRs #370/#374). This issue is the standalone owner of the empirical half. Activation trigger revised in decision-log entry 89 (/brainstorm blocked tickets, 2026-09-08).Why deferred
Nothing is blocked on the evaluation. The policy (risk-tier stance,
authoring.methodprovenance marker, split reviewer qualification — all in spec023) gives reviewers a boundary today. The corpus + controlled run + measurement report is large work whose shape depends on there being enough real LLM-assisted publishes to measure.Activation condition
Start this work when either:
authoring.method: "llm-assisted"— a one-line count overcapabilities/**/contract.jsononmain. Progress: ~3 / 5 as of 2026-09-08 (audio.calibration-plan-create@1.0.0,inference.evidence-normalize@1.0.0,inference.evidence-normalize@1.0.1).#369was the first.023's low-risk deterministic tier fromexperimentaltosupported(this evaluation's data is the precondition for that move).Activation means "begin building the corpus and methodology" — not "there is already enough data for a credible report". The report's own credibility bar (sample size, per-bucket
n) is set by the methodology when it is written, the same split#365uses.Scope when activated (absorbs every unfinished #355 DoD item)
023FR-003 provides most of this chain).experimental→supportedpromotion of the low-risk tier (a later standalone owner decision, per023FR-004 / Governing Relationship).(#355 DoD items "policy by risk tier", "non-Rust vs qualified review guidance", and "no bypass of existing gates" are already delivered in spec
023FR-004 / FR-005 / FR-006 and are not repeated here.)Definition of done
docs/with the five metrics above, sample size, and limitations.docs/decision-log.mdentry recording whether the low-risk tier is promoted tosupported, with the data behind it.023amended (or a successor spec) if the evidence changes the tier stance.Related
023+ this issue)023-authoring-assurance