Skip to content
Merged
Show file tree
Hide file tree
Changes from all commits
Commits
File filter

Filter by extension

Filter by extension

Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
84 changes: 61 additions & 23 deletions CHANGELOG.md
Original file line number Diff line number Diff line change
Expand Up @@ -22,6 +22,46 @@ and this project adheres to [Semantic Versioning](https://semver.org/spec/v2.0.0
placeholder `reasoning_content`: DeepSeek requires it on tool requests and
ignores it otherwise.

### Changed

- `evalshift doctor`'s `evalshift.yaml` row now ends a failure with "— run
`evalshift validate` for details". The row has room for the summary only
("1 schema problem found"), which names neither the offending key nor the
fix; `validate` prints both.

### Removed

- **The top-level `slices:` key is gone from `evalshift.yaml`, and its
removal is breaking.** The block (`name`, `filter`, `applies_to`) was
validated and recorded in the run bundle, but analysis never read it:
slices have always come from example `tags`, one per distinct tag plus
`all`, so none of its fields renamed, filtered or scoped anything, and a
run reports the same slices without it. **Migration: delete the block.**
Per-slice budgets, the one thing it looked like it configured, go under
`migration_policy.slices`, keyed by tag.

A config that still sets `slices:` now **fails to load** instead of being
quietly ignored, the same way `thresholds` has since 1.1.0, naming what
happened and the fix:

```text
`slices` was removed: it never had any effect. Slices come from example `tags` automatically (one per distinct tag, plus `all`). Delete it from evalshift.yaml; per-slice budgets go under migration_policy.slices, keyed by tag.
```

`evalshift validate` prints it; `evalshift doctor` fails its config row
and points at `validate`. `version:` stays `1`, under the same config
version policy. The two shipped example configs that set the block no
longer do, and `tests/unit/test_docs_currency.py` now fails if any doc or
example shows a top-level `slices:` or `thresholds:` key.

Hosted baselines are unaffected for every config that never set the key
(or set `slices: []`): the bundle's `evaluator_config` still carries
`"slices": []`, so `eval_config_hash` is byte-identical to what earlier
CLIs computed and existing baselines keep matching. Deleting a *non-empty*
block does change that hash, so runs pushed afterwards are not comparable
to baselines pushed before until the base branch pushes a run with the
edited config.

### Fixed

- Under SQLAlchemy 2.1, which fresh installs resolve (`sqlalchemy>=2.0`), a
Expand Down Expand Up @@ -87,29 +127,27 @@ and this project adheres to [Semantic Versioning](https://semver.org/spec/v2.0.0
places, and all of them now say what the CLI does. Nothing about CLI
behaviour changed. The one that matters most concerns slices: the docs
said a top-level `slices:` block picks examples by tag, names the slice,
and scopes it to prompts. It does none of that. The block is validated
(the reserved `overall` name is still rejected) and recorded in the run
bundle, but analysis never reads it. Slices come from example `tags`, one
per distinct tag plus `all`, and per-slice budgets are keyed by tag under
`migration_policy.slices`. The docs now say the block is not applied, and
so does the `SliceConfig` docstring, which described `filter` as a Python
expression. Imported agent traces (`traces import`) stay local; the bundle
carries only the replay's own tool-call trace, without tool results or
`model_call` events. The response cache serves only tool-less examples,
so every `run` of an agent suite is live and full price. `--resume`
hashes the suite's path, not its contents. `push <run-id>` uploads an
existing bundle as-is instead of rebuilding it, and `bundle` needs a git
SHA. `--policy-gate` also fails when no `migration_policy` is configured.
The `init` profile table had the wrong `model-upgrade` numbers and no
tool-divergence column. The multi-turn suite example failed to load
because it had no `tools`. The failure-label list was missing
`TOOL_GROUND_TRUTH_MISS`. Upstream model-call failures and evaluator
failures are handled the same way, as errored rows excluded from the
statistics. The GitHub Action docs gained `require-policy` and the other
missing inputs. `record_model_call` examples now pass the required
`tools=`. DOCS.md's header said version 1.0.1; a new check in
`tests/unit/test_docs_currency.py` keeps the version in DOCS.md and
llms-full.txt equal to the package's.
and scopes it to prompts. It did none of that. The block was validated and
recorded in the run bundle, but analysis never read it. Slices come from
example `tags`, one per distinct tag plus `all`, and per-slice budgets are
keyed by tag under `migration_policy.slices`. The docs now say so, and the
block itself is removed in this release (see Removed above). Imported
agent traces (`traces import`) stay local; the bundle carries only the
replay's own tool-call trace, without tool results or `model_call` events.
The response cache serves only tool-less examples, so every `run` of an
agent suite is live and full price. `--resume` hashes the suite's path,
not its contents. `push <run-id>` uploads an existing bundle as-is instead
of rebuilding it, and `bundle` needs a git SHA. `--policy-gate` also fails
when no `migration_policy` is configured. The `init` profile table had the
wrong `model-upgrade` numbers and no tool-divergence column. The
multi-turn suite example failed to load because it had no `tools`. The
failure-label list was missing `TOOL_GROUND_TRUTH_MISS`. Upstream
model-call failures and evaluator failures are handled the same way, as
errored rows excluded from the statistics. The GitHub Action docs gained
`require-policy` and the other missing inputs. `record_model_call`
examples now pass the required `tools=`. DOCS.md's header said version
1.0.1; a new check in `tests/unit/test_docs_currency.py` keeps the version
in DOCS.md and llms-full.txt equal to the package's.

## [1.1.0] - 2026-09-19

Expand Down
27 changes: 17 additions & 10 deletions DOCS.md
Original file line number Diff line number Diff line change
Expand Up @@ -207,7 +207,7 @@ init → doctor → run → evaluate → analy
.json (if policy)
```

- **`doctor`** validates local config and shows which provider keys are visible. Exit 1 only when an existing `evalshift.yaml` fails validation; missing keys are soft warnings. Its second row, `evalshift-sdk`, reports the SDK version the `evalshift` import name resolves to in this environment (`warn` when the SDK is missing or shadowed by an older CLI's leftover files; never a failure). It also reports the toolset each configured suite carries (or the flat `golden.jsonl`) and flags a suite whose examples carry more than one distinct toolset — legal (each example dispatches its own), but also the shape a wiring mistake takes. When a workflow under `.github/workflows/` uses the GitHub Action it adds a `ci pin` row: `ok` (`pinned to <v>`) when CI installs this CLI version, `warn` when the pin is older, absent, or newer than the local CLI (see [Pin drift](#pin-drift)). When the config wires an `llm_judge` evaluator and names both `defaults.source_model` and `target_model`, a `judge family` row warns for every `judge_model` that resolves to the same provider as an arm (self-preference bias; `ok` "from a third family" otherwise, no row when either arm is unset) — advisory, never a failure; `validate` prints the same line and the report repeats it above the verdict (see [`evaluators.llm_judge`](docs/configuration.md#evaluatorsllm_judge)).
- **`doctor`** validates local config and shows which provider keys are visible. Exit 1 only when an existing `evalshift.yaml` fails validation — the row gives the problem count and points at `evalshift validate`, which prints each problem; missing keys are soft warnings. Its second row, `evalshift-sdk`, reports the SDK version the `evalshift` import name resolves to in this environment (`warn` when the SDK is missing or shadowed by an older CLI's leftover files; never a failure). It also reports the toolset each configured suite carries (or the flat `golden.jsonl`) and flags a suite whose examples carry more than one distinct toolset — legal (each example dispatches its own), but also the shape a wiring mistake takes. When a workflow under `.github/workflows/` uses the GitHub Action it adds a `ci pin` row: `ok` (`pinned to <v>`) when CI installs this CLI version, `warn` when the pin is older, absent, or newer than the local CLI (see [Pin drift](#pin-drift)). When the config wires an `llm_judge` evaluator and names both `defaults.source_model` and `target_model`, a `judge family` row warns for every `judge_model` that resolves to the same provider as an arm (self-preference bias; `ok` "from a third family" otherwise, no row when either arm is unset) — advisory, never a failure; `validate` prints the same line and the report repeats it above the verdict (see [`evaluators.llm_judge`](docs/configuration.md#evaluatorsllm_judge)).
- **`run`** parses prompts, validates every example against every prompt, estimates cost, then dispatches `(prompt × example × {source, target})` calls through an async orchestrator under a concurrency semaphore. Responses to tool-less examples are cached (tool-calling examples are always dispatched live — see [Response cache](#response-cache)); progress is checkpointed every 50 completions.
- **`evaluate`** scores each (source, target) pair with the configured evaluators, one `EvalRecord` per pair × evaluator. Scoring runs under the same `defaults.concurrency` semaphore as `run`, and the embedding/judge calls it makes go through the same response cache.
- **`analyze`** runs paired statistics per `(prompt, evaluator, slice)`, applies Benjamini–Hochberg FDR correction, classifies severities, and — when a `migration_policy` is configured — computes a pass/fail verdict.
Expand Down Expand Up @@ -314,7 +314,6 @@ suites: {}
| `prompts` | list, required, ≥1 | Prompt definitions (unique ids enforced) |
| `defaults` | block | Run defaults, below |
| `evaluators` | block | Evaluator configs, see [Evaluators](#evaluators) |
| `slices` | list | Validated and recorded in the bundle (and in its `eval_config_hash`), but **not applied** — slices come from example tags. See below |
| `migration_policy` | block \| absent | Regression budgets, see [Migration policy](#migration-policy-and-ci-gating) |
| `suites` | map | Named suites (`{name: {source: captured\|jsonl, path: ..., evaluators: ..., managed: true}}`); the block between the `>>> evalshift suites` markers is managed by `capture sync`. See [Per-suite evaluators](#per-suite-evaluators) |
| `retention` | block | `max_runs_per_suite` (default 20, `0` disables), `run_ttl_days` (default off) |
Expand All @@ -327,6 +326,14 @@ suites: {}

Nothing replaced it. Delete the block; express any gate you meant by it as a `migration_policy` budget.

A top-level `slices` list was removed the same way. It was validated and recorded in the run bundle, but analysis never read it — slices come from example tags (see [Slices](#slices)) — so a config that still carries it fails to load too:

```text
`slices` was removed: it never had any effect. Slices come from example `tags` automatically (one per distinct tag, plus `all`). Delete it from evalshift.yaml; per-slice budgets go under migration_policy.slices, keyed by tag.
```

Delete the block; a run reports the same slices without it. Hosted baselines are unaffected for any config that never set it (or set `slices: []`): the bundle's evaluator config still carries an empty `slices` list, so `eval_config_hash` does not move. Deleting a *non-empty* block does change that hash: runs pushed afterwards are not comparable to baselines pushed before, until the base branch pushes a run with the edited config.

### `defaults`

| Field | Default | Meaning |
Expand All @@ -341,18 +348,18 @@ Nothing replaced it. Delete the block; express any gate you meant by it as a `mi
| `max_tokens` | `4096` | Completion cap per call (per-prompt `prompts[].max_tokens` overrides). Truncated calls are excluded from the regression statistics |
| `samples_per_example` | `1` (1–20) | Repeats each (prompt, example) this many times per model. Each sample pair is scored on its own; `scores.jsonl` keeps one row per example holding the mean over samples, with the per-sample scores and `delta_variance` under `metadata.samples`. Paired tests run over examples, so `n` is unchanged. Cost and calls multiply by it; the cache keys on the sample index so every sample is a live call |

### `slices`
### Slices

Slices come from the suite, not from this block: every distinct example `tag` becomes a slice under its own name, alongside the implicit `all` slice. Every configured evaluator is analysed once overall and once per slice. Per-slice budgets go under [`migration_policy.slices`](#migration-policy-and-ci-gating), keyed by the tag.
Slices come from the suite; there is nothing to configure. Every distinct example `tag` becomes a slice under its own name, alongside the implicit `all` slice, and every configured evaluator is analysed once overall and once per slice. Per-slice budgets go under [`migration_policy.slices`](#migration-policy-and-ci-gating), keyed by the tag:

```yaml
slices: # validated and recorded in the run bundle; NOT applied by analysis
- name: security
filter: security # a literal tag
applies_to: ["*"] # glob list of prompt ids
migration_policy:
slices:
security: # the example tag
max_overall_regression_rate: 0.0
```

The top-level `slices:` block still loads — it is validated and copied into the run bundle's evaluator config — but analysis does not read it today: `name`, `filter` and `applies_to` rename, filter and scope nothing, and a run reports the same slices with or without it. It is, however, part of the bundle's `eval_config_hash`, so editing or removing it breaks hosted baseline compatibility with earlier runs. `overall` is reserved — it names the run-level scope in the run bundle — and is rejected as a slice `name`, as an example tag, and as a `migration_policy.slices` key.
There is no top-level `slices:` key — it was removed (see [Top-level fields](#top-level-fields)). `overall` is reserved — it names the run-level scope in the run bundle — and is rejected as an example tag and as a `migration_policy.slices` key.

Slices holding exactly the same examples are collapsed to one before any test runs — duplicates restate the same numbers as if they were independent findings and skew the Benjamini–Hochberg correction anti-conservatively (extra copies of a p-value shrink every adjusted p-value in the family, so results look more significant than they are). `all` and any slice named under `migration_policy.slices` always survive; otherwise the provenance tag `captured` (written by `capture promote`) loses to an ordinary tag, then alphabetical order decides. Drops are reported on the terminal and as `collapsed_slices` in `analysis.json`. See [docs/methodology.md](docs/methodology.md).

Expand Down Expand Up @@ -881,7 +888,7 @@ EvalShift follows [Semantic Versioning](https://semver.org). From **1.0.0** onwa
- The HTML report's markup, styling, and internal structure. Its *content* is described here; its DOM is not.
- Console output wording, progress rendering, and log formatting.

**Config schema evolution** has its own rule, and it is deliberately not tied to the CLI's major version: `version:` in `evalshift.yaml` bumps only when a field is renamed, removed, or given a different meaning. Additive fields ride the CLI version instead. See [Config version policy](docs/configuration.md#config-version-policy) for what that requires of your CI pin.
**Config schema evolution** has its own rule, and it is deliberately not tied to the CLI's major version: `version:` in `evalshift.yaml` bumps only when a field is renamed or given a different meaning. Additive fields ride the CLI version instead, and so do removals that fail the load with a message naming the key. See [Config version policy](docs/configuration.md#config-version-policy) for what that requires of your CI pin.

**Renames keep the old name.** When a command is renamed, the previous name stays registered as a hidden alias that still works — it stops being advertised, not accepted. `evalshift all` became `evalshift compare` in 1.0.0 and `all` still runs. Removing such an alias would itself be a breaking change, so it cannot happen inside a major version.

Expand Down
8 changes: 5 additions & 3 deletions README.md
Original file line number Diff line number Diff line change
Expand Up @@ -70,9 +70,11 @@ CI enforces a 90% coverage floor on the source. The CLI is published on PyPI as
`evalshift`, the capture SDK as `evalshift-sdk`, and the hosted service runs at
`api.evalshift.dev`.

The `evalshift.yaml` schema is versioned: `version: 1` changes only for a
breaking change — a field renamed, removed, or given new semantics — so configs
and CI pipelines keep working across releases.
The `evalshift.yaml` schema is versioned: `version: 1` changes only when a
field is renamed or given new semantics. A removed field does not bump it: a
config that still sets one fails to load with an error naming the key, so no
config is ever silently misread across releases. See
[Config version policy](docs/configuration.md#config-version-policy).

## Install

Expand Down
Loading
Loading