Skip to content
Merged
Show file tree
Hide file tree
Changes from all commits
Commits
File filter

Filter by extension

Filter by extension

Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
66 changes: 52 additions & 14 deletions CHANGELOG.md
Original file line number Diff line number Diff line change
Expand Up @@ -29,6 +29,28 @@ and this project adheres to [Semantic Versioning](https://semver.org/spec/v2.0.0
("1 schema problem found"), which names neither the offending key nor the
fix; `validate` prints both.

- Repeat runs of agent suites are now served from the response cache. Every
example that offers tools used to bypass the cache, a leftover from v0.2
when a cache entry could not hold a parsed tool trace, so each `run` of an
agent suite paid for every call again, one per replayed round. Each replayed
round is now its own entry. It is keyed on the canonical model, the prompt
and inputs, the exact message list that round sends (history, current turn,
and the recorded rounds and fixture results fed back), the tool list exactly
as sent and in order (including `strict`), `generation_config` (so
`tool_choice` and `parallel_tool_calls`), the effective temperature and
`max_tokens`, the round index and the sample index. A hit restores the
parsed trace, tokens, cost, latency and finish reason, so the `raw.jsonl`
row is identical to the live one apart from `cached`, which is true only
when every round hit, and the new `cached_rounds` count. A row with any
round served from the cache carries latency from an earlier run, so it stays
out of the report's live latency figures and its latency delta is marked not
comparable, in `report.json` and in the bundle alike. Errors are never
cached: the next run re-sends a failed round and serves the rounds before it
from the cache. Truncated responses are cached and stay flagged, and
`defaults.cache: false` still sends everything live. Existing
`~/.evalshift/cache.db` files keep working: the new `trace_json` column is
added in place on open, and every cached text response still hits.

### Removed

- **The top-level `slices:` key is gone from `evalshift.yaml`, and its
Expand Down Expand Up @@ -72,6 +94,21 @@ and this project adheres to [Semantic Versioning](https://semver.org/spec/v2.0.0
`docs/github-action.md`, and `docs/hosted.md` now name `run:create` +
`run:read` + `policy:read` for the CI key.

- A run whose target was served entirely from the cache while the source ran
live showed a -100% latency change in the HTML report header and in the
run insights. Latency averages cover only calls measured live on this run,
so a role with none of them averages 0, which means unmeasured, not
instant. Both now say the latency change is not comparable unless both
roles measured some latency live.

- Two `evalshift` processes opening the response cache at the same moment
(parallel CI jobs, or a `run` beside an `evaluate`) could crash one of them
with `table cached_calls already exists` when the cache file was new. Both
had checked for the table, and the slower one then tried to create it too.
Opening the cache now treats that error, and its migration counterpart
`duplicate column name`, as "another process already did it" and carries
on; any other schema error is still raised.

- Under SQLAlchemy 2.1, which fresh installs resolve (`sqlalchemy>=2.0`), a
`CacheStore` opened on an in-memory SQLite database could silently lose
concurrent writes. The default on-disk cache used by CLI runs was not
Expand Down Expand Up @@ -142,20 +179,21 @@ and this project adheres to [Semantic Versioning](https://semver.org/spec/v2.0.0
block itself is removed in this release (see Removed above). Imported
agent traces (`traces import`) stay local; the bundle carries only the
replay's own tool-call trace, without tool results or `model_call` events.
The response cache serves only tool-less examples, so every `run` of an
agent suite is live and full price. `--resume` hashes the suite's path,
not its contents. `push <run-id>` uploads an existing bundle as-is instead
of rebuilding it, and `bundle` needs a git SHA. `--policy-gate` also fails
when no `migration_policy` is configured. The `init` profile table had the
wrong `model-upgrade` numbers and no tool-divergence column. The
multi-turn suite example failed to load because it had no `tools`. The
failure-label list was missing `TOOL_GROUND_TRUTH_MISS`. Upstream
model-call failures and evaluator failures are handled the same way, as
errored rows excluded from the statistics. The GitHub Action docs gained
`require-policy` and the other missing inputs. `record_model_call`
examples now pass the required `tools=`. DOCS.md's header said version
1.0.1; a new check in `tests/unit/test_docs_currency.py` keeps the version
in DOCS.md and llms-full.txt equal to the package's.
The response cache served only tool-less examples, so every `run` of an
agent suite was live and full price (it now serves them too; see Changed).
`--resume` hashes the suite's path, not its contents. `push <run-id>`
uploads an existing bundle as-is instead of rebuilding it, and `bundle`
needs a git SHA. `--policy-gate` also fails when no `migration_policy` is
configured. The `init` profile table had the wrong `model-upgrade` numbers
and no tool-divergence column. The multi-turn suite example failed to load
because it had no `tools`. The failure-label list was missing
`TOOL_GROUND_TRUTH_MISS`. Upstream model-call failures and evaluator
failures are handled the same way, as errored rows excluded from the
statistics. The GitHub Action docs gained `require-policy` and the other
missing inputs. `record_model_call` examples now pass the required
`tools=`. DOCS.md's header said version 1.0.1; a new check in
`tests/unit/test_docs_currency.py` keeps the version in DOCS.md and
llms-full.txt equal to the package's.

## [1.1.0] - 2026-09-19

Expand Down
8 changes: 4 additions & 4 deletions DOCS.md
Original file line number Diff line number Diff line change
Expand Up @@ -208,7 +208,7 @@ init → doctor → run → evaluate → analy
```

- **`doctor`** validates local config and shows which provider keys are visible. Exit 1 only when an existing `evalshift.yaml` fails validation — the row gives the problem count and points at `evalshift validate`, which prints each problem; missing keys are soft warnings. Its second row, `evalshift-sdk`, reports the SDK version the `evalshift` import name resolves to in this environment (`warn` when the SDK is missing or shadowed by an older CLI's leftover files; never a failure). It also reports the toolset each configured suite carries (or the flat `golden.jsonl`) and flags a suite whose examples carry more than one distinct toolset — legal (each example dispatches its own), but also the shape a wiring mistake takes. When a workflow under `.github/workflows/` uses the GitHub Action it adds a `ci pin` row: `ok` (`pinned to <v>`) when CI installs this CLI version, `warn` when the pin is older, absent, or newer than the local CLI (see [Pin drift](#pin-drift)). When the config wires an `llm_judge` evaluator and names both `defaults.source_model` and `target_model`, a `judge family` row warns for every `judge_model` that resolves to the same provider as an arm (self-preference bias; `ok` "from a third family" otherwise, no row when either arm is unset) — advisory, never a failure; `validate` prints the same line and the report repeats it above the verdict (see [`evaluators.llm_judge`](docs/configuration.md#evaluatorsllm_judge)).
- **`run`** parses prompts, validates every example against every prompt, estimates cost, then dispatches `(prompt × example × {source, target})` calls through an async orchestrator under a concurrency semaphore. Responses to tool-less examples are cached (tool-calling examples are always dispatched live — see [Response cache](#response-cache)); progress is checkpointed every 50 completions.
- **`run`** parses prompts, validates every example against every prompt, estimates cost, then dispatches `(prompt × example × {source, target})` calls through an async orchestrator under a concurrency semaphore. Responses are cached — per call, and per replayed round for tool-calling examples (see [Response cache](#response-cache)); progress is checkpointed every 50 completions.
- **`evaluate`** scores each (source, target) pair with the configured evaluators, one `EvalRecord` per pair × evaluator. Scoring runs under the same `defaults.concurrency` semaphore as `run`, and the embedding/judge calls it makes go through the same response cache.
- **`analyze`** runs paired statistics per `(prompt, evaluator, slice)`, applies Benjamini–Hochberg FDR correction, classifies severities, and — when a `migration_policy` is configured — computes a pass/fail verdict.
- **`report`** renders the single-file HTML report (no external assets; works offline and attaches cleanly to a PR or email), and writes the machine-written [run insights](#run-insights) narrative unless `--no-insights` is passed. The page opens on a verdict / advisory-signal / economics panel row and a six-cell run strip (examples, calls, failed-or-truncated, spend, latency Δ, mean score Δ), then the executive summary, the narrative, one section per prompt, and the methodology. Every figure on it is derived from the run's own artefacts; the deltas in the header are the run-level rollup of the per-prompt economics. Top regressions are collapsed cards — expand one for the trace diff, the tool diffs and the conversation context. The report is dark-only.
Expand Down Expand Up @@ -237,7 +237,7 @@ Run ids look like `r_20260722_golden_a1b2c3` (`r_<date>_<suite-slug>_<hex>`). Un

### Response cache

Live responses to **tool-less** examples are cached in SQLite at `~/.evalshift/cache.db`, keyed by SHA-256 over canonical JSON of `(model, prompt, inputs, temperature, max_tokens[, history])` — plus `generation_config`, the toolset fingerprint, the round index and the sample index when each is set — with a 7-day TTL. Re-running an identical tool-less evaluation costs no run-stage calls. **Examples with a non-empty toolset are not cached:** every `run` of an agent suite dispatches them live, at full price. The evaluate-stage embedding and judge caches below still apply to them. Disable per-project with `defaults.cache: false`; wipe with `evalshift cache clear`.
Live run-stage responses are cached in SQLite at `~/.evalshift/cache.db`, keyed by SHA-256 over canonical JSON of `(model, prompt, inputs, temperature, max_tokens[, history])` — plus `generation_config`, the tools array as sent, the round index and the sample index when each is set — with a 7-day TTL. Re-running an identical evaluation costs no run-stage calls, agent suites included. An example with a non-empty toolset gets one entry per replayed round, keyed on the canonical model, the prompt and inputs, **the exact message list that round sends** (history, current turn, and the recorded rounds and fixture results teacher forcing feeds back), the tool list exactly as sent, in order (tool names, descriptions, schemas and `strict` — reordering the tools is a miss), `generation_config` (so `tool_choice` and `parallel_tool_calls`), the effective temperature and `max_tokens`, the round index and the sample index — editing a fixture re-sends only the rounds after it. A hit restores the parsed tool trace (calls with their ids and arguments, final text, refusal), tokens, cost, latency and finish reason, so the `raw.jsonl` row matches the live one apart from `cached` and `cached_rounds`; a multi-round example counts as cached only when every round was a hit. A row with any round served from the cache (`cached_rounds > 0`) carries latency recorded on an earlier run, so it is left out of the report's live latency figures and its latency delta is marked not comparable. Errors are never cached (the next run re-sends a failed round; the rounds before it come from the cache), truncated responses are cached and stay flagged, and a hit records the original call's cost and latency. A `cache.db` written by an earlier version is upgraded in place on open and keeps its entries; several `evalshift` processes may open it at once. Disable per-project with `defaults.cache: false`; wipe with `evalshift cache clear`.

The cache covers the evaluate stage too: `semantic` embeddings are keyed by `(embedding model, text)`, and `llm_judge` verdicts by `(judge model, criterion, source output, target output)`. The judge key uses a canonical A/B ordering, so the per-call orientation randomization doesn't halve the hit rate — the orientation that was actually used is recorded with the verdict and replayed on a hit, leaving `metadata.target_was_a` faithful.

Expand Down Expand Up @@ -918,11 +918,11 @@ Keys are consumed by LiteLLM at call time; EvalShift itself never stores or tran

### Will `run` cost me money?

Yes — every `run` calls a real model. Before dispatch you get a worst-case cost estimate (assumes every completion hits the registry `default_max_tokens`, 4096 — actual cost is usually much lower); above $10 it asks for confirmation. The cache makes repeat runs of unchanged tool-less calls free; tool-calling examples are dispatched live (full price) on every run. Cheapest iteration loop: small suite first, cache on.
Yes — every `run` calls a real model. Before dispatch you get a worst-case cost estimate (assumes every completion hits the registry `default_max_tokens`, 4096 — actual cost is usually much lower); above $10 it asks for confirmation. The cache makes repeat runs of unchanged calls free, tool-calling examples included. Cheapest iteration loop: small suite first, cache on.

### A model call failed mid-run

The error is recorded on that call in `raw.jsonl`; the run completes. At evaluate time the affected pair gets an errored row (a 0.5/0.5 placeholder with the error attached) that is excluded from the statistics, so it can't masquerade as a regression or an improvement. Re-running the same command re-uses cached successes (tool-less examples only — tool-calling examples are all dispatched again) and retries the failures (errored calls in a *resumed* run are not retried — start a fresh run to retry them).
The error is recorded on that call in `raw.jsonl`; the run completes. At evaluate time the affected pair gets an errored row (a 0.5/0.5 placeholder with the error attached) that is excluded from the statistics, so it can't masquerade as a regression or an improvement. Re-running the same command re-uses cached successes (for a multi-round tool example, the rounds before the failed one) and retries the failures (errored calls in a *resumed* run are not retried — start a fresh run to retry them).

### `--resume` aborts with a config-hash mismatch

Expand Down
10 changes: 9 additions & 1 deletion docs/agents.md
Original file line number Diff line number Diff line change
Expand Up @@ -82,7 +82,6 @@ defaults:
source_model: gemini-2.5-flash
target_model: gemini-3.1-flash-lite-preview
judge_model: gemini-3.1-pro-preview
cache: false

evaluators:
# Top-level: what every suite is scored with. Tool evaluators do not belong
Expand Down Expand Up @@ -223,6 +222,15 @@ the results recorded in the same round; a name match never reaches across a
example as diverged if **any** replayed round diverged.
* The cost estimate counts one call per replayed round; the progress bar still
counts examples.
* Each round is its own response-cache entry, keyed on exactly what that round
sends: the prompt, the recorded rounds and fixture results fed back, the
tool list exactly as sent and in order (including `strict`),
`generation_config` and the round index. A repeat run is served from the
cache; editing round *k*'s fixtures re-sends only the rounds after it, and a
round that errored is re-sent while the rounds before it are not. The
example's row counts as cached only when every round was a hit; a partly
cached row's latency is left out of the live latency figures, since some of
it was measured on an earlier run.

`expected_tools` is `expected_tool_rounds[0]` under both settings. `--rounds
all` no longer flattens every round into `expected_tools` — that yardstick was
Expand Down
2 changes: 1 addition & 1 deletion docs/configuration.md
Original file line number Diff line number Diff line change
Expand Up @@ -366,7 +366,7 @@ A list of prompt definitions. Each entry has:
| `judge_model` | string | `gemini-3.1-flash-lite-preview` | Default LLM-as-judge model. |
| `insights_model`| string | (none) | Model that writes the run-insights narrative rendered in `report.html` and uploaded with the bundle. Falls back to `judge_model` when unset — writing analytical prose is a harder task than a pairwise A/B verdict, so it is worth tuning separately. See [Run insights](#run-insights). |
| `concurrency` | int | 10 (1 ≤ x ≤ 64) | Max in-flight LLM calls during `evalshift run` **and** `evalshift evaluate` (the embedding and judge calls made while scoring). |
| `cache` | bool | `true` | Read/write the local SQLite cache at `~/.evalshift/cache.db`. Covers run-stage completions of tool-less examples (examples that offer tools are always dispatched live) plus `semantic` embeddings and `llm_judge` verdicts. |
| `cache` | bool | `true` | Read/write the local SQLite cache at `~/.evalshift/cache.db`. Covers run-stage completions — for examples that offer tools, one entry per replayed round — plus `semantic` embeddings and `llm_judge` verdicts. See [Response cache](https://github.com/babaliauskas/evalshift-cli/blob/main/DOCS.md#response-cache) for the key. |
| `max_cost_usd` | float | 50.0 | Soft ceiling reserved for future enforcement. The pre-flight cost prompt currently triggers above $10 (skip with `--yes`). |
| `max_tokens` | int | 4096 (`> 0`) | Completion length cap sent to every model call. Raise it if outputs are being truncated (the provider returns `finish_reason == "length"`); a `prompts[].max_tokens` entry overrides it per prompt. Truncated calls are detected, surfaced in the report, and **excluded from the regression statistics** so a cut-off output can't manufacture a false regression. |
| `samples_per_example` | int | 1 (1 ≤ x ≤ 20) | How many times each `(prompt, example)` is sent to **each** model. Above 1, every sample is its own live call (the cache keys on the sample index), sample *i* of the source is scored against sample *i* of the target, and the example's row in `scores.jsonl` becomes the **mean over samples** with the per-sample scores and the within-example `delta_variance` under `metadata.samples`. The paired tests still run over examples, not samples, so this reduces noise without inflating `n`. Cost and the call count multiply by it; only worth turning on for a model that samples non-deterministically (see the report banner). See [Methodology](methodology.md#limitations-to-be-aware-of). |
Expand Down
5 changes: 2 additions & 3 deletions docs/evaluators.md
Original file line number Diff line number Diff line change
Expand Up @@ -230,7 +230,6 @@ A 100-example suite with 1 prompt and 4 evaluators (2 structural +
* Evaluate: 200 embedding calls + 100 judge calls

LiteLLM's pricing data drives the pre-flight estimate; the local
SQLite cache absorbs identical re-runs of tool-less examples, evaluate-stage
embedding and judge calls included. Examples that offer tools are dispatched
live on every run; only their evaluate-stage calls are cached. Evaluate dispatches its calls under
SQLite cache absorbs identical re-runs, examples that offer tools and
evaluate-stage embedding and judge calls included. Evaluate dispatches its calls under
`defaults.concurrency`, same as the run stage.
7 changes: 3 additions & 4 deletions docs/faq.md
Original file line number Diff line number Diff line change
Expand Up @@ -137,10 +137,9 @@ you see `≤ $0.17` and the run actually cost $0.03, that's expected.
## How do I lower the cost of a run?

* **Set the SQLite cache to be on** (it's the default). A re-run of
the exact same configuration makes no run-stage calls for tool-less
examples. Examples that offer tools are not cached: an agent suite's
run stage is live, at full price, every time. Evaluate-stage embedding
and judge calls are cached either way.
the exact same configuration makes no run-stage calls, agent suites
included: examples that offer tools are cached one entry per replayed
round. Evaluate-stage embedding and judge calls are cached too.
* **Use cheaper models.** The model registry assigns sensible
defaults but you can drop everything to flash/mini/haiku tier.
* **Skip the LLM judge.** Structural and semantic evaluators are
Expand Down
Loading
Loading