This repository was archived by the owner on Sep 15, 2026. It is now read-only.
Effective cost metric — slice 1 (run-level) - #66
Merged
Merged
Conversation
effective_cost = realized_cost / realized_quality. Slice 1 run-level (build now: gate verdict + per-task cost list -> effective-cost.json + report section). Slice 2 per-worker learned leaderboard (whole-run, single-producer-weighted). Routing re-rank named + deferred. Paper-informed (Price Reversal Phenomenon, arXiv 2603.23971): cost numerator is a labelled cost-tier estimate behind a -CostResolver seam; attempts first-class for the future multi-turn multiplier. Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
… record Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
There was a problem hiding this comment.
Pull request overview
Implements slice 1 of the run-level “effective cost” metric (effective_cost = realized_cost / realized_quality) by introducing a pure PowerShell library, wiring it into the Conductor’s run finalization to emit effective-cost.json + a ## Effective cost report section when a gate verdict exists, and adding tests + docs + bootstrap deployment updates.
Changes:
- Add
scripts/effective-cost-lib.ps1with pure functions to compute quality scalar, run cost, effective cost, worker breakdown, record shape, and report formatting. - Wire the metric into
scripts/conductor-lib.ps1(Complete-Run+ DAG walk accumulation) and extend conductor tests to assert the new artifact/section behavior. - Update bootstrap deployment/test assertions and bump plugin version; add spec + implementation plan docs.
Reviewed changes
Copilot reviewed 9 out of 9 changed files in this pull request and generated 3 comments.
Show a summary per file
| File | Description |
|---|---|
| scripts/effective-cost-lib.ps1 | New pure library implementing the effective-cost math + record/section formatting. |
| scripts/conductor-lib.ps1 | Wires effective-cost generation into Complete-Run and accumulates per-task costs during the DAG walk. |
| scripts/test-effective-cost-lib.ps1 | New hermetic test suite covering library behavior (bands, cost, breakdown, record, section). |
| scripts/test-conductor-lib.ps1 | Adds coverage for the new run artifact + report section emission behavior. |
| scripts/bootstrap.ps1 | Adds effective-cost-lib.ps1 to the deployed script manifest. |
| scripts/test-bootstrap.ps1 | Asserts bootstrap dry-run output includes effective-cost-lib.ps1. |
| docs/superpowers/specs/2026-06-22-effective-cost-metric-design.md | New design spec for the metric and slice 1 wiring. |
| docs/superpowers/plans/2026-06-22-effective-cost-metric-slice1.md | New implementation plan for slice 1. |
| .claude-plugin/plugin.json | Version bump to 1.4.0-rc.3. |
| Add-RunDecision -RunDir $RunDir -Decision $dec | ||
| [void]$decisions.Add($dec) | ||
| } | ||
| [void]$taskCosts.Add(@{ id = $task.id; worker = ([string]$r.chose); cost = $tspend }) |
Comment on lines
+59
to
+63
| - **Banded by verdict** (monotonic — an accept always outranks a polish, a | ||
| polish always outranks a reject): | ||
| - `accept` → `[0.7, 1.0]` | ||
| - `polish` → `[0.3, 0.7)` | ||
| - `reject` → `(0.0, 0.3]` |
| ## Global Constraints | ||
|
|
||
| - **Pure layer is I/O-free** — `effective-cost-lib.ps1` does no file/network/dispatch work; all inputs arrive as parameters (mirrors `saturation-lib.ps1`). | ||
| - **Quality scalar is banded, refined, floored** — `accept ∈ [0.7,1.0]`, `polish ∈ [0.3,0.7)`, `reject ∈ (0,0.3]`; refined within band by counts (weights `critical=0.5`, `important=0.2`, `minor=0.05`); global floor `0.05`. Never returns ≤ 0. |
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to subscribe to this conversation on GitHub.
Already have an account?
Sign in.
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Slice 1 of the quality-adjusted effective-cost metric (v1.4 line, plugin 1.4.0-rc.3).
effective_cost = realized_cost / realized_quality— what you actually paid for the quality you actually got. Lower is better. Run-level: when the d058 acceptance gate produces a verdict, the Conductor writes a 6th run artifacteffective-cost.json+ a## Effective costreport section. No gate → byte-for-byte unchanged.effective-cost-lib.ps1(pure):Get-QualityScalar(banded by verdict, refined by counts, floored >0),Get-RunCost(labelledestimatebasis +-CostResolverseam +attemptsfield),Get-EffectiveCost,Get-WorkerBreakdown,New-EffectiveCostRecord,Format-EffectiveCostSection.conductor-lib.ps1: DAG walk accumulates a per-task cost list;Complete-Rungains-TaskCosts, emits the record + section on a gate verdict, returnseffective_cost.test-effective-cost-lib.ps1(E1–E33),test-conductor-lib.ps1(T80–T86),test-bootstrap.ps1assert. All green.Final adversarial opus review: READY TO MERGE, 0 Critical / 0 Important (3 Minor deferred — polish-band boundary tie, unused
$lo, dead-but-defensive guard).Paper-informed (Price Reversal Phenomenon, arXiv 2603.23971): cost numerator is honestly a cost-tier estimate behind the
-CostResolverseam.Spec: docs/superpowers/specs/2026-06-22-effective-cost-metric-design.md
Plan: docs/superpowers/plans/2026-06-22-effective-cost-metric-slice1.md
🤖 Generated with Claude Code