Skip to content
This repository was archived by the owner on Sep 15, 2026. It is now read-only.
Merged
Show file tree
Hide file tree
Changes from all commits
Commits
File filter

Filter by extension

Filter by extension

Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
2 changes: 1 addition & 1 deletion .claude-plugin/plugin.json
Original file line number Diff line number Diff line change
@@ -1,7 +1,7 @@
{
"name": "baton",
"displayName": "Baton",
"version": "1.4.0",
"version": "1.4.1-rc.1",
"description": "Pass the baton. Conduct the fleet. Claude Code as command-and-control for a fleet of coding LLMs — capability routing, cost engine, jobs, decisions, and a knowledge base.",
"author": { "name": "Kevin Rank", "url": "https://github.com/Ryfter" },
"repository": "https://github.com/Ryfter/baton",
Expand Down
11 changes: 9 additions & 2 deletions commands/effective-cost.md
Original file line number Diff line number Diff line change
Expand Up @@ -16,8 +16,9 @@ cheap on paper but cost more once rejects and polish rounds are counted.
Whole-run, single-producer-weighted attribution: a worker that did one task in a
mixed run is credited by its cost share, not condemned for the rest. `confidence`
rises with run count and the fraction of clean single-producer runs; low-confidence
rows are flagged **tentative**. Advisory / legibility only — it never changes
routing.
rows are flagged **tentative**. This command is advisory / read-only. Routing only
consumes the leaderboard when you opt in with `learned_routing: true` in the
box-private `~/.baton/fleet.yaml` (slice 3, d060 — see step 4); off by default.

## Steps

Expand All @@ -38,3 +39,9 @@ routing.
3. Summarize in plain language: who the cheapest-quality-adjusted worker is, which
rows are still tentative (and why — too few clean runs), and any worker that
looks cheap by tier but ranks poorly once quality is folded in.

4. Routing consumer (opt-in): with `learned_routing: true` in the box-private
`~/.baton/fleet.yaml`, `/baton:go` routing biases worker selection in economy
mode by this same leaderboard (slice 3, d060) — a learned-expensive worker yields
toward the next cost tier, bounded to an adjacent-tier shift and confidence-gated.
Off by default → routing is unchanged and this command stays purely informational.
388 changes: 388 additions & 0 deletions docs/superpowers/plans/2026-06-26-effective-cost-rerank.md

Large diffs are not rendered by default.

202 changes: 202 additions & 0 deletions docs/superpowers/specs/2026-06-26-effective-cost-rerank-design.md
Original file line number Diff line number Diff line change
@@ -0,0 +1,202 @@
# Confidence-Gated Learned-Cost Re-rank — Design (effective-cost slice 3)

> **Status:** approved design (2026-06-26). Decision: **d060**. Next: implementation plan.
> **Slice line:** v1.4 post-MVP, slice 3. Builds on slice 1 (`effective-cost.json`
> per-run records) and slice 2 (`Get-WorkerEffectiveCost` leaderboard fold).
> **Realizes:** the named-deferred §4.3 of `2026-06-22-effective-cost-metric-design.md`
> — the d026 "router that learns" payoff. Listed price is a lie; learned effective
> cost corrects it (Price Reversal Phenomenon, arXiv 2603.23971).

## 1. Problem

Slices 1–2 measure and surface effective cost but are **advisory only** —
`Select-Capability` still ranks by the ordinal cost tier (`local` < `free` <
`paid`) and never consults what a worker has *actually* cost per unit quality. A
cheap-tier worker that the Acceptance Gate keeps rejecting is genuinely more
expensive than a mid-tier worker that clears — but routing can't see that yet.
Slice 3 closes the loop: the learned `eff_cost_mean` biases the economy ranking,
**once the signal is trusted (confidence-gated) and only when explicitly enabled.**

## 2. Scope guard (what this is NOT)

- **Not** default-on. A global `learned_routing: true` opt-in gates the entire
mechanism. Default off → `Select-Capability` is **byte-for-byte unchanged**, the
same invariant the saturation driver holds.
- **Not** a tier-ordinal replacement. The cost tier remains the budget guardrail.
The learned signal **shifts** a worker's effective rank by a bounded amount; it
never lets a worker leap two tiers (d060: bounded adjacent-tier).
- **Not** a champion-mode change. Champion ("just the best quality", cost-blind)
ignores the learned cost signal, exactly as it ignores saturation.
- **Not** new I/O on the metric. It folds the same box-private
`effective-cost.json` records slice 2 already reads. No new artifact.

## 3. Design

### 3.1 Pure decision function (`effective-cost-lib.ps1`)

```
Get-LearnedCostAdjustment
-Worker <string>
-Board <object[]> # rows from Get-WorkerEffectiveCost
[-MinConfidence <double> = 0.5]
[-MaxShift <double> = 1.0]
-> @{ adjust = <double>; confidence = <double>; reason = <string|null> }
```

I/O-free, like `Get-SaturationDecision`. Rules:

1. **Inert when untrusted or absent.** Find the worker's row in `Board`. If absent,
or its `confidence < MinConfidence` → `@{ adjust = 0.0; confidence = <row or 0>;
reason = $null }`. Untrusted signal never moves routing.
2. **Baseline = median over trusted PEERS (excluding the evaluated worker).**
Compute the median `eff_cost_mean` across rows whose `confidence >= MinConfidence`
**and `worker != Worker`** — a relative re-rank compares a candidate against its
*alternatives*, not a pool containing itself (self-inclusion would dampen the
signal and make confidence-weighting unobservable for a worker at the median).
Untrusted rows neither move nor anchor. If fewer than 1 trusted peer, or median
`<= 0` → inert (fail-open; a single-trusted-worker fleet produces no bias).
3. **Signed, bounded, symmetric shift.** `logr = ln(eff_cost_mean / median)` — `0`
at median, **positive when the worker is worse** (more expensive per quality),
negative when better. Clamp to `[-MaxShift, MaxShift]`.
4. **Confidence-weighted.** Scale the clamped shift by
`w = clamp((confidence - MinConfidence) / (1 - MinConfidence), 0, 1)` — a
just-cleared-the-bar worker barely moves; a fully-confident worker gets the full
bounded shift. `adjust = clampedShift * w`, rounded to 4 dp.
5. **Reason** (non-null only when `adjust != 0`): `"learned eff_cost <m> vs fleet
median <med> (conf <c>) -> <+/-adjust> tier"`.

Positive `adjust` = worse = **up-rank toward a more expensive tier** (yields).
Negative `adjust` = better = **down-rank toward a cheaper tier** (preferred).

### 3.2 Effective-rank helper (`saturation-lib.ps1`, beside `Get-EffectiveTierRank`)

```
Get-LearnedTierRank -CostTier <string> [-Saturating <bool>] [-Adjust <double>]
-> <double>
```

- **Saturation wins.** `if ($Saturating) { return -1 }` — a worker spending its free
allotment is the strongest down-rank; learned bias does not fight it.
- Else `rank = (Get-CostTierRank $CostTier) + $Adjust`; **floored at -1** so the
learned signal can never undercut saturation's −1. Returns a `double` (fractional
ranks separate same-tier workers by learned cost — within-tier ordering falls out
for free).

### 3.3 Wiring (`routing-lib.ps1` → `Select-Capability`, economy branch only)

`Select-Capability` gains an injectable `-RunsRoot` parameter (default
`(Join-Path (Get-BatonHome) 'runs')`), mirroring its existing `-RatingsPath` /
`-JournalPath` / `-UsagePath` seams so the board source is overridable in tests and
never touches real `~/.baton`. After the §3b saturation block, **when learned
routing is enabled and a leaderboard exists**, annotate each surviving candidate
with its adjustment:

```powershell
# 3c. Learned-cost re-rank (d060) — opt-in, economy-only, confidence-gated.
$learnedOn = Get-LearnedRoutingEnabled -FleetPath $FleetPath # global switch, default $false
$board = @()
if ($learnedOn -and $SelectionMode -eq 'economy') {
$records = Read-EffectiveCostRecords -RunsRoot $RunsRoot # injectable seam, default (Get-BatonHome)/runs
# NOTE: no @() around the call. Get-WorkerEffectiveCost returns ,@($rows) (unary-comma),
# which direct assignment unwraps to the rows array; wrapping it in @() would re-nest it
# into a 1-element array holding the rows. The @($board).Count guard below re-wraps safely.
if (@($records).Count -gt 0) { $board = Get-WorkerEffectiveCost -Records $records }
}
foreach ($c in $filtered) {
$c | Add-Member -NotePropertyName learned_adjust -NotePropertyValue 0.0 -Force
if ($learnedOn -and $SelectionMode -eq 'economy' -and @($board).Count -gt 0) {
$d = Get-LearnedCostAdjustment -Worker $c.name -Board $board
$c.learned_adjust = [double]$d.adjust
if ($d.reason) { $c.why = "$($c.why); $($d.reason)" }
}
}
```

The economy sort's **primary key** changes from `Get-EffectiveTierRank` to
`Get-LearnedTierRank`:

```powershell
@{e={ Get-LearnedTierRank $_.cost_tier ([bool]$_.saturate) ([double]$_.learned_adjust) }}, `
@{e={ if ([bool]$_.saturate) { [double]$_.sat_util } else { 0 } }}, `
@{e={ -$_.quality }}, @{e='name'}
```

When `learned_routing` is off (default), `learned_adjust` is `0.0` for every
candidate and `Get-LearnedTierRank … 0` ≡ `Get-EffectiveTierRank` → **identical
ranking, byte-for-byte**. Champion branch untouched.

### 3.4 Config switch (`Get-LearnedRoutingEnabled`)

A top-level `learned_routing: true` in `fleet.yaml` (box-private) enables it.
Helper reads the fleet file, returns `$true` only for a literal boolean `$true`
(same strict-opt-in coercion the saturation driver uses for non-canonical YAML
false tokens). Absent / false / non-boolean → `$false`.

`Read-EffectiveCostRecords` is the same record-reader `fleet-effective-cost.ps1`
already defines (globs `*/effective-cost.json`, try/catch skips malformed,
`return ,@($records)`); slice 3 lifts it into `effective-cost-lib.ps1` so both the
CLI and routing share one implementation (DRY).

## 4. Data flow

```
effective-cost.json (per run, box-private) ── Read-EffectiveCostRecords
│
▼
Get-WorkerEffectiveCost (fold) ──► leaderboard rows
│
▼
Get-LearnedCostAdjustment (per candidate, confidence-gated, bounded ±MaxShift)
│
▼
Get-LearnedTierRank (saturation-floored) ──► Select-Capability economy sort key
```

## 5. Error handling & invariants

- **Default-off byte-for-byte:** `learned_routing` unset → no record read, every
`learned_adjust = 0.0`, ranking identical to pre-slice-3. Verified in a test that
ranks the same fleet with the switch off and asserts order is unchanged.
- **Fail-open:** absent/empty/malformed records → empty board → inert (no throw).
Median over zero trusted rows → inert.
- **No divide-by-zero / no NaN:** median `<= 0` → inert; `eff_cost_mean > 0` by
construction (quality floored `> 0` in slice 1). `ln` only ever sees a positive
ratio.
- **Saturation supremacy:** `Get-LearnedTierRank` returns `-1` when saturating
regardless of `Adjust`; the floor keeps learned bias `>= -1` otherwise.
- **Bounded reach:** `|adjust| <= MaxShift` (default 1.0) → at most an adjacent-tier
shift; a 2-tier leap is impossible.
- **Box-private:** the leaderboard is folded from `$BATON_HOME/runs/…` and never
leaves the box. `references/fleet.yaml` (shared) carries only the field doc for
`learned_routing`; no box values.
- **Champion unchanged:** the champion branch never reads the board.
- **PowerShell house rules:** no param/local named `$args`/`$input`/`$event`/
`$matches`/`$host`; parenthesize function calls inside comparisons; guard
unary-comma flatten on empty (`return @()` for the empty case); files
`utf8NoBOM`.

## 6. Testing (hermetic — temp dirs, injected board, zero network, never real ~/.baton)

- **`test-effective-cost-lib.ps1`** (extend): `Get-LearnedCostAdjustment` —
worse-than-median → positive adjust; cheaper → negative; below `MinConfidence` →
0; absent worker → 0; bound respected (`|adjust| <= MaxShift`) even for an
extreme ratio; confidence-weighting (just-cleared worker moves less than a
fully-confident one at the same ratio); empty/single-row board → 0.
- **`test-saturation-lib.ps1`** (extend): `Get-LearnedTierRank` — saturating → -1
ignoring Adjust; non-saturating local `+0.5` → 0.5; floor at -1 for a large
negative Adjust; `Adjust 0` ≡ `Get-EffectiveTierRank`.
- **`test-routing-lib.ps1`** (extend): switch **off** → ranking identical to a
captured baseline (byte-for-byte invariant); switch **on** with a seeded board
where a learned-bad cheap worker yields to a learned-good neighbour; champion mode
ignores the board; `Get-LearnedRoutingEnabled` strict-opt-in (true/false/absent/
non-boolean token).
- **`test-bootstrap.ps1`**: manifest already deploys `effective-cost-lib.ps1`,
`saturation-lib.ps1`, `routing-lib.ps1` — no manifest change; assert unchanged.
- **Plugin:** `.claude-plugin/plugin.json` → `1.4.1-rc.1` (post-1.4.0 line).

## 7. Decision

- **d060** — learned effective-cost re-rank uses a **bounded adjacent-tier
adjustment** (±`MaxShift`, default 1.0), confidence-gated (`MinConfidence` 0.5),
default-off, economy-only, saturation-floored. Alternatives (within-tier-only;
full effective rank) rejected — see the record.
9 changes: 9 additions & 0 deletions references/fleet.yaml
Original file line number Diff line number Diff line change
Expand Up @@ -68,6 +68,15 @@ capability_floors:
# propose culling them, the router can never claim them (they're unregistered).
keep_list: ['*heretic*']

# learned_routing (optional, top-level) — BOX-PRIVATE, default OFF. When set true in
# your live ~/.baton/fleet.yaml, Select-Capability biases its ECONOMY ranking by each
# worker's LEARNED effective cost (slice 3, d060): a worker that has cost more per unit
# quality (Acceptance-Gate verdicts ÷ spend, folded from $BATON_HOME/runs) yields toward
# the next cost tier; a learned-cheap worker reaches toward the tier below. Bounded to an
# adjacent-tier shift, confidence-gated (only trusted rows move routing), saturation-
# floored. Champion mode ignores it. Off / absent => routing is byte-for-byte unchanged.
# learned_routing: true # set ONLY in your live box file, never in this shared seed.

providers:
- name: claude-cli
kind: cli
Expand Down
78 changes: 78 additions & 0 deletions scripts/effective-cost-lib.ps1
Original file line number Diff line number Diff line change
Expand Up @@ -206,6 +206,84 @@ function Get-WorkerEffectiveCost {
return ,@($sorted)
}

function Read-EffectiveCostRecords {
<# Read effective-cost.json records, parse each, skip malformed. Two source modes:
-Glob an explicit path/glob to the record FILES (e.g. 'D:/runs/*/effective-cost.json');
resolved with Get-ChildItem -Path. This is the CLI's `--runs` surface.
-RunsRoot a runs ROOT directory; recurse for 'effective-cost.json' under it. Routing's seam.
-Glob wins when both are given. Neither supplied / missing path -> empty array.
Pure/dependency-free: all paths are parameters. #>
param(
[string]$RunsRoot,
[string]$Glob
)
$files = if (-not [string]::IsNullOrWhiteSpace($Glob)) {
@(Get-ChildItem -Path $Glob -File -ErrorAction SilentlyContinue)
} elseif ((-not [string]::IsNullOrWhiteSpace($RunsRoot)) -and (Test-Path $RunsRoot)) {
@(Get-ChildItem -Path $RunsRoot -Filter 'effective-cost.json' -Recurse -File -ErrorAction SilentlyContinue)
} else { @() }
Comment on lines +220 to +224
if ($files.Count -eq 0) { return @() }
$records = foreach ($f in $files) {
try { Get-Content -LiteralPath $f.FullName -Raw | ConvertFrom-Json } catch { continue }
}
$records = @($records)
if ($records.Count -eq 0) { return @() }
return ,@($records)
}

function Get-LearnedRoutingEnabled {
<# Read the fleet YAML for the top-level learned_routing switch.
$true ONLY for a literal boolean true; absent/false/non-boolean -> $false.
Pure/dependency-free: path is a parameter. #>
param([Parameter(Mandatory)][string]$FleetPath)
if (-not (Test-Path $FleetPath)) { return $false }
foreach ($line in (Get-Content -LiteralPath $FleetPath)) {
if ($line -match '^\s*learned_routing\s*:\s*(.+?)\s*$') {
$val = $Matches[1].Trim().Trim('"').Trim("'")
return ($val -eq 'true')
}
}
Comment on lines +240 to +245
return $false
}

function Get-LearnedCostAdjustment {
<# Map a worker's learned eff_cost_mean vs the trusted-fleet median into a
bounded, confidence-weighted rank shift. Positive = worse (yields up a tier);
negative = better (preferred). Inert when untrusted/absent. Pure. #>
param(
[Parameter(Mandatory)][string]$Worker,
[object[]]$Board = @(),
[double]$MinConfidence = 0.5,
[double]$MaxShift = 1.0
)
$rows = @($Board)
$me = $rows | Where-Object { [string]$_.worker -eq $Worker } | Select-Object -First 1
$trusted = @($rows | Where-Object { [double]$_.confidence -ge $MinConfidence -and [double]$_.eff_cost_mean -gt 0 -and [string]$_.worker -ne $Worker })
$conf = if ($me) { [double]$me.confidence } else { 0.0 }
if (-not $me -or $conf -lt $MinConfidence -or [double]$me.eff_cost_mean -le 0 -or $trusted.Count -lt 1) {
return @{ adjust = 0.0; confidence = $conf; reason = $null }
}
$vals = @($trusted | ForEach-Object { [double]$_.eff_cost_mean } | Sort-Object)
$mid = [int][math]::Floor($vals.Count / 2)
$median = if ($vals.Count % 2 -eq 1) { $vals[$mid] } else { ($vals[$mid - 1] + $vals[$mid]) / 2.0 }
if ($median -le 0) { return @{ adjust = 0.0; confidence = $conf; reason = $null } }
$logr = [math]::Log(([double]$me.eff_cost_mean / $median))
$clamped = [math]::Max(-$MaxShift, [math]::Min($MaxShift, $logr))
# Guard the degenerate band MinConfidence = 1.0: the denominator is 0, and 0/0 = NaN
# slips past the [-0,1] clamp below (NaN comparisons are always false). A worker that
# clears a bar of 1.0 has conf = 1.0 exactly -> full weight.
$denom = 1.0 - $MinConfidence
$w = if ($denom -le 0) { 1.0 } else { ($conf - $MinConfidence) / $denom }
if ($w -lt 0) { $w = 0.0 } elseif ($w -gt 1) { $w = 1.0 }
$adjust = [math]::Round(($clamped * $w), 4)
$reason = $null
if ($adjust -ne 0) {
$sign = if ($adjust -gt 0) { '+' } else { '' }
$reason = "learned eff_cost $('{0:0.00}' -f [double]$me.eff_cost_mean) vs fleet median $('{0:0.00}' -f $median) (conf $('{0:0.00}' -f $conf)) -> $sign$adjust tier"
}
return @{ adjust = $adjust; confidence = $conf; reason = $reason }
}

function Format-EffectiveCostLeaderboard {
<# Render a Get-WorkerEffectiveCost leaderboard as a plain-text report block.
Rows arrive already cheapest-first (do not re-sort). Low-confidence rows
Expand Down
16 changes: 2 additions & 14 deletions scripts/fleet-effective-cost.ps1
Original file line number Diff line number Diff line change
Expand Up @@ -19,22 +19,10 @@ param(
$ErrorActionPreference = 'Stop'
. (Join-Path $PSScriptRoot 'effective-cost-lib.ps1')

function Read-EffectiveCostRecords {
<# Glob effective-cost.json records and parse them; a malformed file is
skipped, never fatal. Returns an array (possibly empty). #>
param([string]$Root, [string]$Glob)
$pattern = if ($Glob) { $Glob } else { Join-Path $Root '*/effective-cost.json' }
$records = @()
foreach ($f in @(Get-ChildItem -Path $pattern -File -ErrorAction SilentlyContinue)) {
try { $records += (Get-Content -LiteralPath $f.FullName -Raw | ConvertFrom-Json) }
catch { [Console]::Error.WriteLine("skipped unreadable record: $($f.FullName)") }
}
return ,@($records)
}

switch ($Subcommand) {
'report' {
$records = Read-EffectiveCostRecords -Root $RunsRoot -Glob $Runs
# --runs (a path/glob to record files) overrides the default $BATON_HOME runs root.
$records = if ($Runs) { Read-EffectiveCostRecords -Glob $Runs } else { Read-EffectiveCostRecords -RunsRoot $RunsRoot }
$board = Get-WorkerEffectiveCost -Records @($records) -MinConfidenceRuns $MinConfidenceRuns
if ($Json) {
# -InputObject (not pipe): a piped array unrolls, so ConvertTo-Json
Expand Down
Loading