Skip to content

Empirical evaluation of LLM-assisted capability authoring (deferred child of #355) #371

Description

@enricopiovesan

Deferred from #355 via /brainstorm 341 355 365 (decision-log entry 85); #355 is now closed — its policy half shipped as approved, CI-enforced spec 023-authoring-assurance (PRs #370/#374). This issue is the standalone owner of the empirical half. Activation trigger revised in decision-log entry 89 (/brainstorm blocked tickets, 2026-09-08).

Why deferred

Nothing is blocked on the evaluation. The policy (risk-tier stance, authoring.method provenance marker, split reviewer qualification — all in spec 023) gives reviewers a boundary today. The corpus + controlled run + measurement report is large work whose shape depends on there being enough real LLM-assisted publishes to measure.

Activation condition

Start this work when either:

  1. ≥ 5 published capabilities carry authoring.method: "llm-assisted" — a one-line count over capabilities/**/contract.json on main. Progress: ~3 / 5 as of 2026-09-08 (audio.calibration-plan-create@1.0.0, inference.evidence-normalize@1.0.0, inference.evidence-normalize@1.0.1). #369 was the first.
  2. or there is an owner decision to move 023's low-risk deterministic tier from experimental to supported (this evaluation's data is the precondition for that move).

Activation means "begin building the corpus and methodology" — not "there is already enough data for a credible report". The report's own credibility bar (sample size, per-bucket n) is set by the methodology when it is written, the same split #365 uses.

Scope when activated (absorbs every unfinished #355 DoD item)

  • A versioned evaluation corpus covering deterministic/pure, validation, effectful, and explicitly-excluded high-risk capability classes.
  • Every generated candidate linked to its contract, source revision, artifact digest, test evidence, and human review decision (spec 023 FR-003 provides most of this chain).
  • Report measures: first-pass validation rate, security/policy rejection rate, reviewer rework, defect escape, time-to-accepted artifact. Methodology and limitations published.
  • Contract-derived tests + at least one property-based or metamorphic check where the contract admits them.
  • The report's data sets the numeric bar for the experimentalsupported promotion of the low-risk tier (a later standalone owner decision, per 023 FR-004 / Governing Relationship).

(#355 DoD items "policy by risk tier", "non-Rust vs qualified review guidance", and "no bypass of existing gates" are already delivered in spec 023 FR-004 / FR-005 / FR-006 and are not repeated here.)

Definition of done

  • Evaluation corpus published (private/controlled per Establish evidence and policy for LLM-assisted capability authoring #355 D2; publishable aggregate methodology + results).
  • Report published under docs/ with the five metrics above, sample size, and limitations.
  • A docs/decision-log.md entry recording whether the low-risk tier is promoted to supported, with the data behind it.
  • 023 amended (or a successor spec) if the evidence changes the tier stance.

Related

Activity

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Labels

    enhancementNew feature or requestqualityQuality gates, CI, coverage, and engineering standards

    Type

    No type

    Projects

    No projects

      Milestone

      No milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions