Skip to content
Merged
Show file tree
Hide file tree
Changes from all commits
Commits
File filter

Filter by extension

Filter by extension


Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
19 changes: 19 additions & 0 deletions CHANGELOG.md
Original file line number Diff line number Diff line change
Expand Up @@ -7,6 +7,21 @@ and this project adheres to [Semantic Versioning](https://semver.org/spec/v2.0.0

## [Unreleased]

### Added

- DeepSeek is a supported provider. `deepseek-flash` and `deepseek-v4-pro`
are in the model registry; bare `deepseek-*` ids (as recorded by a capture
of an app calling `api.deepseek.com` through the OpenAI client) resolve to
`deepseek/…`; `DEEPSEEK_API_KEY` is checked before a run and shown by
`evalshift doctor`; tool-call evals parse DeepSeek responses; a DeepSeek
judge grading a DeepSeek arm gets the judge-family warning; and
`evalshift init --provider deepseek` scaffolds a DeepSeek project. DeepSeek's
default thinking mode ignores `temperature`, so DeepSeek arms and a DeepSeek
judge carry the report's non-determinism banner, and every assistant turn
replayed from the recording (tool rounds and chat history) is sent with a
placeholder `reasoning_content`: DeepSeek requires it on tool requests and
ignores it otherwise.

### Fixed

- Under SQLAlchemy 2.1, which fresh installs resolve (`sqlalchemy>=2.0`), a
Expand All @@ -22,6 +37,10 @@ and this project adheres to [Semantic Versioning](https://semver.org/spec/v2.0.0
sessions one at a time on a single-connection pool. This works the same
on SQLAlchemy 2.0 and 2.1.

- `capture sync` priced calls recorded under a bare id at $0 when LiteLLM
keys the model under both spellings but can only price the provider-prefixed
one (DeepSeek). The price lookup now tries the provider-prefixed id first.

- README.md, DOCS.md, llms-full.txt, four `docs/` pages, and AGENTS.md
advertised `evalshift all --push` as the command to run. `all` has been a
hidden alias for `compare` since 1.0.0 — it still works, and always will —
Expand Down
42 changes: 39 additions & 3 deletions DOCS.md
Original file line number Diff line number Diff line change
Expand Up @@ -75,6 +75,7 @@ API keys go in the environment, never in config:
export GEMINI_API_KEY=... # or GOOGLE_API_KEY
export OPENAI_API_KEY=...
export ANTHROPIC_API_KEY=...
export DEEPSEEK_API_KEY=...
```

Only the providers your configured models use need a key. `evalshift doctor` shows which keys are visible.
Expand Down Expand Up @@ -133,7 +134,7 @@ See [Project setup](#project-setup) and [Capturing from production](#capturing-f

`init` options:

- `--provider gemini|openai|anthropic` — which provider's model ids the scaffold uses (prompted on a TTY; defaults to `gemini` otherwise). Gemini and OpenAI scaffolds include an embedding-based semantic evaluator; the Anthropic scaffold comments it out (no embedding endpoint).
- `--provider gemini|openai|anthropic|deepseek` — which provider's model ids the scaffold uses (prompted on a TTY; defaults to `gemini` otherwise). Gemini and OpenAI scaffolds include an embedding-based semantic evaluator; the Anthropic and DeepSeek scaffolds comment it out (no embedding endpoint).
- `--profile` — pre-tuned migration-policy budgets:

| Profile | regression ≤ | critical ≤ | equivalence ≥ | arg drift ≤ | cost Δ ≤ | latency Δ ≤ |
Expand Down Expand Up @@ -382,7 +383,7 @@ Derivation, on every `capture sync`: no row offered a toolset → **no block** (

### Model ids

Model resolution is deliberately permissive: a small built-in registry maps aliases to canonical `provider/model` ids, and anything unknown is passed through with provider inferred from the prefix (`gemini-*` → Google, `claude-*` → Anthropic, `gpt-*`/`o1-*`/`o3-*` → OpenAI). LiteLLM is the call-time authority — **any model LiteLLM supports works**; the registry never gates. Before a live run the CLI checks that the inferred provider's API key env var is set.
Model resolution is deliberately permissive: a small built-in registry maps aliases to canonical `provider/model` ids, and anything unknown is passed through with provider inferred from the prefix (`gemini-*` → Google, `claude-*` → Anthropic, `gpt-*`/`o1-*`/`o3-*` → OpenAI, `deepseek-*` → DeepSeek). LiteLLM is the call-time authority — **any model LiteLLM supports works**; the registry never gates. Before a live run the CLI checks that the inferred provider's API key env var is set.

---

Expand Down Expand Up @@ -800,7 +801,7 @@ Common conventions: `-c/--config` defaults to `./evalshift.yaml`; run artefacts
### Pipeline

**`evalshift init`** — scaffold a minimal capture-first `evalshift.yaml`.
`-f/--force` · `-d/--directory <dir>` · `--ci` · `--wire-agents/--no-wire-agents` (default on) · `--provider gemini|openai|anthropic` · `--profile model-upgrade|cost-reduction|local-model|quantization|provider-switch` (default `model-upgrade`)
`-f/--force` · `-d/--directory <dir>` · `--ci` · `--wire-agents/--no-wire-agents` (default on) · `--provider gemini|openai|anthropic|deepseek` · `--profile model-upgrade|cost-reduction|local-model|quantization|provider-switch` (default `model-upgrade`)
Without `--ci`, warns after writing when an existing workflow under `.github/workflows/` pins an older CLI than this one, or none at all (see [Pin drift](#pin-drift)); `init --ci` writes the pin itself and does not warn about the file it just wrote.

**`evalshift doctor`** — environment/config check. Exit 1 only on an invalid existing config. Row 2, `evalshift-sdk`, confirms `import evalshift` is the SDK (`warn` when missing or shadowed, never a failure). Reports the toolset each configured suite carries and flags a suite whose examples carry more than one distinct toolset. The suite-side checks cover every suite in the config's `suites:` block, falling back to `./golden.jsonl` when none are wired. Adds a `ci pin` row when a workflow uses the GitHub Action (`warn` on pin drift, never a failure) and a `judge family` row when an `llm_judge` judge shares a provider with a configured arm (`warn`, never a failure).
Expand Down Expand Up @@ -888,6 +889,7 @@ EvalShift follows [Semantic Versioning](https://semver.org). From **1.0.0** onwa
| `GEMINI_API_KEY` / `GOOGLE_API_KEY` | — | Google auth (either works) |
| `OPENAI_API_KEY` | — | OpenAI auth (also the default semantic embedding model) |
| `ANTHROPIC_API_KEY` | — | Anthropic auth |
| `DEEPSEEK_API_KEY` | — | DeepSeek auth |
| `EVALSHIFT_NONINTERACTIVE` | unset | Non-empty → skip the cost-confirmation prompt (implied `--yes`); set in scaffolded CI |
| `EVALSHIFT_MAX_RUNS` | unset | Override `retention.max_runs_per_suite`; `0`/`none`/`unlimited`/`off` disables count pruning |
| `EVALSHIFT_DIR` | `.evalshift` | Base dir for SDK captures the `capture` commands read |
Expand Down Expand Up @@ -930,6 +932,40 @@ Yes — one example per turn with a recorded `history` prefix, replayed teacher-

Anything LiteLLM supports. The built-in registry only provides aliases and metadata; unknown ids pass through with provider inferred from the id prefix. Verify a model with `evalshift test-call -m <id>`.

### Does EvalShift work with DeepSeek?

Yes. Export `DEEPSEEK_API_KEY` and use DeepSeek's API ids, `deepseek-flash`
or `deepseek-v4-pro`. A bare `deepseek-*` id (what a capture records when your
app calls `api.deepseek.com` through the OpenAI client) gets the `deepseek/`
prefix automatically. `evalshift init --provider deepseek` scaffolds a
DeepSeek project. Three things differ from other providers:

- **Sampling is not controlled.** Both models run in thinking mode by
default, which accepts `temperature` and ignores it. EvalShift keeps thinking
on, because that is what your application runs, so DeepSeek arms are marked
non-deterministic in the report. A DeepSeek judge is marked
non-deterministic too. Raise `defaults.samples_per_example` when the verdict
matters.
- **Replayed assistant turns carry an empty reasoning chain.** Every assistant
turn replayed from the recording, tool rounds and chat history alike, is
sent with the single-space `reasoning_content` placeholder the API accepts.
The recording holds no DeepSeek reasoning to pass back. DeepSeek requires
the field on any request with tools, where an empty chain may degrade
multi-turn answer quality, and ignores it otherwise.
- **No embeddings.** DeepSeek has no embedding endpoint. The `semantic`
evaluator needs an OpenAI or Gemini embedding model and its key, which is
why the DeepSeek scaffold ships it commented out.

DeepSeek served by another host (self-hosted open weights, or a cloud region
of your choice) goes through that host's LiteLLM prefix (`hosted_vllm/`,
`azure_ai/`, `bedrock/`, ...) and its environment variables. Tool calls parse
the same way, but the key pre-check and the notes above apply to the
`deepseek/` API only. LiteLLM also reads `DEEPSEEK_API_BASE` to point the
`deepseek/` provider at a DeepSeek-compatible endpoint. A local Ollama model
named like `deepseek-r1` needs its prefix, `ollama/deepseek-r1`, when you name
it as a run arm; a capture that recorded the bare name is treated as the
DeepSeek API, and its estimated capture cost uses DeepSeek's API price.

### Do I need LangChain / a specific framework?

No. EvalShift compares model behaviour, not framework code: prompts come from config or AST-parsed source, each example's toolset from a capture-derived sidecar or an inline `tools:` list, ground truth from captures. Framework-side timelines can be scored via [external traces](#external-agent-traces).
Expand Down
2 changes: 1 addition & 1 deletion README.md
Original file line number Diff line number Diff line change
Expand Up @@ -301,7 +301,7 @@ assets, works offline) has:
Your prompts and suite stay local for `doctor`, `run`, `evaluate`, `analyze`,
and `report`. The only outbound calls in local mode are to the LLM providers you
configure — any provider LiteLLM supports, called with your own API keys.
Anthropic, OpenAI and Google ids additionally get a curated pricing and
Anthropic, DeepSeek, Google and OpenAI ids additionally get a curated pricing and
capability entry; everything else is passed through with the provider inferred
from the id.

Expand Down
2 changes: 1 addition & 1 deletion docs/configuration.md
Original file line number Diff line number Diff line change
Expand Up @@ -593,7 +593,7 @@ row and `evalshift validate` a matching `⚠` line — never a failure, because
API key. The report repeats the note above the verdict whenever a judge that
actually contributed `llm_judge` rows shares a family with an arm, and
`report.json` carries it as `judge_family_overlap`. "Family" is the provider
the model id resolves to (`anthropic`, `openai`, `google`); ids the registry
the model id resolves to (`anthropic`, `deepseek`, `google`, `openai`); ids the registry
cannot place never match.

The judge sees both outputs (with random A/B order to defang positional
Expand Down
38 changes: 36 additions & 2 deletions docs/faq.md
Original file line number Diff line number Diff line change
Expand Up @@ -48,13 +48,47 @@ account for them rather than silently dropping examples.
## What models does EvalShift support?

Anything LiteLLM supports. The `evalshift_cli.models.registry` provides
friendly aliases and sane defaults for common models (Claude, GPT,
Gemini), but **the registry is advisory, not gating**. A model id
friendly aliases and sane defaults for common models (Claude, DeepSeek,
Gemini, GPT), but **the registry is advisory, not gating**. A model id
that isn't in the registry — for example a fresh preview from a
vendor playground — gets passed through to LiteLLM with a
prefix-inferred provider. LiteLLM is the source of truth at call
time.

## Does EvalShift work with DeepSeek?

Yes. Export `DEEPSEEK_API_KEY` and use DeepSeek's API ids, `deepseek-flash`
or `deepseek-v4-pro`. A bare `deepseek-*` id (what a capture records when your
app calls `api.deepseek.com` through the OpenAI client) gets the `deepseek/`
prefix automatically. `evalshift init --provider deepseek` scaffolds a
DeepSeek project. Three things differ from other providers:

- **Sampling is not controlled.** Both models run in thinking mode by
default, which accepts `temperature` and ignores it. EvalShift keeps thinking
on, because that is what your application runs, so DeepSeek arms are marked
non-deterministic in the report. A DeepSeek judge is marked
non-deterministic too. Raise `defaults.samples_per_example` when the verdict
matters.
- **Replayed assistant turns carry an empty reasoning chain.** Every assistant
turn replayed from the recording, tool rounds and chat history alike, is
sent with the single-space `reasoning_content` placeholder the API accepts.
The recording holds no DeepSeek reasoning to pass back. DeepSeek requires
the field on any request with tools, where an empty chain may degrade
multi-turn answer quality, and ignores it otherwise.
- **No embeddings.** DeepSeek has no embedding endpoint. The `semantic`
evaluator needs an OpenAI or Gemini embedding model and its key, which is
why the DeepSeek scaffold ships it commented out.

DeepSeek served by another host (self-hosted open weights, or a cloud region
of your choice) goes through that host's LiteLLM prefix (`hosted_vllm/`,
`azure_ai/`, `bedrock/`, ...) and its environment variables. Tool calls parse
the same way, but the key pre-check and the notes above apply to the
`deepseek/` API only. LiteLLM also reads `DEEPSEEK_API_BASE` to point the
`deepseek/` provider at a DeepSeek-compatible endpoint. A local Ollama model
named like `deepseek-r1` needs its prefix, `ollama/deepseek-r1`, when you name
it as a run arm; a capture that recorded the bare name is treated as the
DeepSeek API, and its estimated capture cost uses DeepSeek's API price.

## Can I resume a run after Ctrl+C / a crash?

Yes. `evalshift run --resume` finds the latest in-progress run for
Expand Down
1 change: 1 addition & 0 deletions docs/getting-started.md
Original file line number Diff line number Diff line change
Expand Up @@ -46,6 +46,7 @@ Set whichever providers you intend to use:
export ANTHROPIC_API_KEY=<anthropic-api-key>
export OPENAI_API_KEY=<openai-api-key>
export GEMINI_API_KEY=<gemini-api-key>
export DEEPSEEK_API_KEY=<deepseek-api-key>
```

## 3. Scaffold your project
Expand Down
4 changes: 2 additions & 2 deletions docs/sdk.md
Original file line number Diff line number Diff line change
Expand Up @@ -73,8 +73,8 @@ intercepted call — sync, async and streaming — records one `model_call` with
the model id, the tools offered, the calls the model **requested**
(`requested_tool_calls`), input, output, token usage, latency and the tool-use
generation settings (`tool_choice`, `parallel_tool_calls`, a tool's `strict`
flag). `wrap_openai` with a `base_url` covers OpenAI-compatible servers (Ollama,
vLLM, Groq, OpenRouter). If you keep `record_model_call`, pass
flag). `wrap_openai` with a `base_url` covers OpenAI-compatible servers
(DeepSeek, Ollama, vLLM, Groq, OpenRouter). If you keep `record_model_call`, pass
`requested_tool_calls=extract_requested_tool_calls(response)` on every model
call: `capture sync` then scores against what the model *asked for* rather than
what the app executed (`promotion_source: requested`), and `run` replays the
Expand Down
Loading
Loading