Skip to content
This repository was archived by the owner on Sep 15, 2026. It is now read-only.

Effective cost metric — slice 1 (run-level) - #66

Merged
Ryfter merged 5 commits into
masterfrom
effective-cost-slice1
Jun 22, 2026
Merged

Ryfter merged 5 commits into
masterfrom
effective-cost-slice1

Conversation

@Ryfter

@Ryfter Ryfter commented Jun 22, 2026

Copy link
Copy Markdown
Owner

Slice 1 of the quality-adjusted effective-cost metric (v1.4 line, plugin 1.4.0-rc.3).

effective_cost = realized_cost / realized_quality — what you actually paid for the quality you actually got. Lower is better. Run-level: when the d058 acceptance gate produces a verdict, the Conductor writes a 6th run artifact effective-cost.json + a ## Effective cost report section. No gate → byte-for-byte unchanged.

  • effective-cost-lib.ps1 (pure): Get-QualityScalar (banded by verdict, refined by counts, floored >0), Get-RunCost (labelled estimate basis + -CostResolver seam + attempts field), Get-EffectiveCost, Get-WorkerBreakdown, New-EffectiveCostRecord, Format-EffectiveCostSection.
  • conductor-lib.ps1: DAG walk accumulates a per-task cost list; Complete-Run gains -TaskCosts, emits the record + section on a gate verdict, returns effective_cost.
  • Tests: test-effective-cost-lib.ps1 (E1–E33), test-conductor-lib.ps1 (T80–T86), test-bootstrap.ps1 assert. All green.

Final adversarial opus review: READY TO MERGE, 0 Critical / 0 Important (3 Minor deferred — polish-band boundary tie, unused $lo, dead-but-defensive guard).

Paper-informed (Price Reversal Phenomenon, arXiv 2603.23971): cost numerator is honestly a cost-tier estimate behind the -CostResolver seam.

Spec: docs/superpowers/specs/2026-06-22-effective-cost-metric-design.md
Plan: docs/superpowers/plans/2026-06-22-effective-cost-metric-slice1.md

🤖 Generated with Claude Code

Ryfter and others added 5 commits June 22, 2026 17:27
effective_cost = realized_cost / realized_quality. Slice 1 run-level
(build now: gate verdict + per-task cost list -> effective-cost.json +
report section). Slice 2 per-worker learned leaderboard (whole-run,
single-producer-weighted). Routing re-rank named + deferred. Paper-informed
(Price Reversal Phenomenon, arXiv 2603.23971): cost numerator is a labelled
cost-tier estimate behind a -CostResolver seam; attempts first-class for the
future multi-turn multiplier.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
… record

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
Copilot AI review requested due to automatic review settings June 22, 2026 23:44
@Ryfter
Ryfter merged commit 633d653 into master Jun 22, 2026

Copilot AI left a comment

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Pull request overview

Implements slice 1 of the run-level “effective cost” metric (effective_cost = realized_cost / realized_quality) by introducing a pure PowerShell library, wiring it into the Conductor’s run finalization to emit effective-cost.json + a ## Effective cost report section when a gate verdict exists, and adding tests + docs + bootstrap deployment updates.

Changes:

  • Add scripts/effective-cost-lib.ps1 with pure functions to compute quality scalar, run cost, effective cost, worker breakdown, record shape, and report formatting.
  • Wire the metric into scripts/conductor-lib.ps1 (Complete-Run + DAG walk accumulation) and extend conductor tests to assert the new artifact/section behavior.
  • Update bootstrap deployment/test assertions and bump plugin version; add spec + implementation plan docs.

Reviewed changes

Copilot reviewed 9 out of 9 changed files in this pull request and generated 3 comments.

Show a summary per file
File Description
scripts/effective-cost-lib.ps1 New pure library implementing the effective-cost math + record/section formatting.
scripts/conductor-lib.ps1 Wires effective-cost generation into Complete-Run and accumulates per-task costs during the DAG walk.
scripts/test-effective-cost-lib.ps1 New hermetic test suite covering library behavior (bands, cost, breakdown, record, section).
scripts/test-conductor-lib.ps1 Adds coverage for the new run artifact + report section emission behavior.
scripts/bootstrap.ps1 Adds effective-cost-lib.ps1 to the deployed script manifest.
scripts/test-bootstrap.ps1 Asserts bootstrap dry-run output includes effective-cost-lib.ps1.
docs/superpowers/specs/2026-06-22-effective-cost-metric-design.md New design spec for the metric and slice 1 wiring.
docs/superpowers/plans/2026-06-22-effective-cost-metric-slice1.md New implementation plan for slice 1.
.claude-plugin/plugin.json Version bump to 1.4.0-rc.3.

Comment thread scripts/conductor-lib.ps1
Add-RunDecision -RunDir $RunDir -Decision $dec
[void]$decisions.Add($dec)
}
[void]$taskCosts.Add(@{ id = $task.id; worker = ([string]$r.chose); cost = $tspend })
Comment on lines +59 to +63
- **Banded by verdict** (monotonic — an accept always outranks a polish, a
polish always outranks a reject):
- `accept` → `[0.7, 1.0]`
- `polish` → `[0.3, 0.7)`
- `reject` → `(0.0, 0.3]`
## Global Constraints

- **Pure layer is I/O-free** — `effective-cost-lib.ps1` does no file/network/dispatch work; all inputs arrive as parameters (mirrors `saturation-lib.ps1`).
- **Quality scalar is banded, refined, floored** — `accept ∈ [0.7,1.0]`, `polish ∈ [0.3,0.7)`, `reject ∈ (0,0.3]`; refined within band by counts (weights `critical=0.5`, `important=0.2`, `minor=0.05`); global floor `0.05`. Never returns ≤ 0.
Sign up for free to subscribe to this conversation on GitHub. Already have an account? Sign in.

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants