Not during local runs. doctor, run, evaluate, analyze, and
report operate locally. Every provider API call goes directly from your
machine to the LLM provider you configured — any provider LiteLLM supports —
using your own API keys.
The local SQLite cache at ~/.evalshift/cache.db only contains
provider responses for your prompts and inputs.
Hosted uploads are explicit. bundle packages the completed
local run artifacts into run_bundle.json.gz without uploading them. push
and compare --push upload that bundle to the hosted backend for your project.
One local stage does call a provider with your data beyond the run itself:
report generates the run-insights narrative, sending the worst regressions'
inputs and outputs to defaults.insights_model. That is your provider, not
EvalShift's, and --no-insights turns it off.
It is the run-insights narrative — plain-language prose written by
defaults.insights_model (falling back to judge_model) explaining the
verdict, the advisory signal, the economics and what changed behaviourally in
the worst regressions.
The prose is machine-written; the numbers are not. Every figure is computed
from the run and handed to the model pre-rendered as a string to copy
verbatim, and any numeric token in the output that was not supplied causes the
generation to be rejected and retried. After two bad generations the CLI ships
deterministic templated prose instead (shown as model none). So a figure in
that block is the same figure as in the tables below it, or it is not there at
all.
It costs one model call per run, is cached in insights.json, and skips
itself when no API key is configured. Disable it with
evalshift report --no-insights or evalshift compare --no-insights.
The orchestrator records the error in raw.jsonl (with error="...")
and moves on. The run still completes. In the evaluation phase the pair
gets an errored row (a 0.5/0.5 placeholder with error set): it stays in
scores.jsonl for inspection and is excluded from the statistics, so a
failed call can't masquerade as a regression or an improvement.
Anything LiteLLM supports. The evalshift_cli.models.registry provides
friendly aliases and sane defaults for common models (Claude, DeepSeek,
Gemini, GPT), but the registry is advisory, not gating. A model id
that isn't in the registry — for example a fresh preview from a
vendor playground — gets passed through to LiteLLM with a
prefix-inferred provider. LiteLLM is the source of truth at call
time.
Yes. Export DEEPSEEK_API_KEY and use DeepSeek's API ids, deepseek-flash
or deepseek-v4-pro. A bare deepseek-* id (what a capture records when your
app calls api.deepseek.com through the OpenAI client) gets the deepseek/
prefix automatically. evalshift init --provider deepseek scaffolds a
DeepSeek project. Three things differ from other providers:
- Sampling is not controlled. Both models run in thinking mode by
default, which accepts
temperatureand ignores it. EvalShift keeps thinking on, because that is what your application runs, so DeepSeek arms are marked non-deterministic in the report. A DeepSeek judge is marked non-deterministic too. Raisedefaults.samples_per_examplewhen the verdict matters. - Replayed assistant turns carry an empty reasoning chain. Every assistant
turn replayed from the recording, tool rounds and chat history alike, is
sent with the single-space
reasoning_contentplaceholder the API accepts. The recording holds no DeepSeek reasoning to pass back. DeepSeek requires the field on any request with tools, where an empty chain may degrade multi-turn answer quality, and ignores it otherwise. - No embeddings. DeepSeek has no embedding endpoint. The
semanticevaluator needs an OpenAI or Gemini embedding model and its key, which is why the DeepSeek scaffold ships it commented out.
DeepSeek served by another host (self-hosted open weights, or a cloud region
of your choice) goes through that host's LiteLLM prefix (hosted_vllm/,
azure_ai/, bedrock/, ...) and its environment variables. Tool calls parse
the same way, but the key pre-check and the notes above apply to the
deepseek/ API only. LiteLLM also reads DEEPSEEK_API_BASE to point the
deepseek/ provider at a DeepSeek-compatible endpoint. A local Ollama model
named like deepseek-r1 needs its prefix, ollama/deepseek-r1, when you name
it as a run arm; a capture that recorded the bare name is treated as the
DeepSeek API, and its estimated capture cost uses DeepSeek's API price.
Yes. evalshift run --resume finds the latest in-progress run for
the project, validates that the config and the suite path haven't changed
since, and continues from where it left off. Already-completed calls
(including ones that errored at the LLM layer) are skipped.
A config change or a different suite path between attempts aborts the resume — start a fresh run instead. The suite's contents are not checked, so after editing examples, start a fresh run yourself.
Sign in through the hosted web app, then approve CLI login in the browser:
evalshift login # defaults to https://api.evalshift.dev
evalshift whoamiSet project: org-slug/project-slug in evalshift.yaml or pass
--project org-slug/project-slug, then run:
evalshift compare --suite-name <suite> --yes --pushSee Hosted EvalShift for credential precedence, bundle contents, and troubleshooting.
Why does the max cost row in evalshift compare look so much higher than the actual Total cost in the report?
The pre-flight figure is a worst-case ceiling, not a forecast.
evalshift compare (and evalshift run) prices each call as if the model
emits its full registry default_max_tokens of completion (4096).
Real completions — especially agent-style runs that produce short
tool-call decisions — are usually far shorter than the cap, so the
actual Total cost in the report typically lands well below the
displayed ceiling.
The figure is conservative on purpose: the cost-confirmation prompt
(triggered above $10) wants to over-warn rather than under-warn. If
you see ≤ $0.17 and the run actually cost $0.03, that's expected.
- Set the SQLite cache to be on (it's the default). A re-run of the exact same configuration makes no run-stage calls, agent suites included: examples that offer tools are cached one entry per replayed round. Evaluate-stage embedding and judge calls are cached too.
- Use cheaper models. The model registry assigns sensible defaults but you can drop everything to flash/mini/haiku tier.
- Skip the LLM judge. Structural and semantic evaluators are
much cheaper. Drop the
evaluators.llm_judgesection to disable the judge entirely. - Cap with
max_cost_usdindefaults(a future tightening will hard-enforce; currently a soft ceiling).
It means the id you passed isn't in EvalShift's curated registry. The id is sent to LiteLLM as-is (with provider prefix inferred from the prefix). If LiteLLM doesn't know the model either, you'll get a clean error from the provider when you make the call.
Two common causes:
- Every delta is identical. When the variance in deltas is near
zero, the test is skipped and severity defaults to
none. - Your sample size is too small. With
n < 5the test is skipped and severity isinsufficient(notnone).
If you expected a real signal, double-check your evaluator output range — many "all the same" cases are evaluators returning a constant.
Three common causes, all by design:
- Every configured evaluator is advisory (
blocking: false— the freshevalshift initstate), so nothing gates quality. Promote evaluators to blocking as your suite grows. The cost and latency budgets still apply — they read the run's calls, so a breach there reportsfail, notinconclusive. - A rate budget was breached but the 95% Wilson interval can't confirm it at this suite size — grow the suite.
- All comparisons were
insufficient(n < 5).
analyze and compare print the specific reason and the recommended fix
under the verdict line, and record them in migration_decision.json
(reason / recommendations).
You don't need LangChain to use EvalShift. Each golden-suite example
carries its own toolset — a toolset_ref pointing at a sidecar, or an
inline tools list (Anthropic-shape or OpenAI-shape, either works). The
usual path is automatic: capture your agent with the evalshift-sdk (works
regardless of framework — LangChain, a manual loop, anything) and
capture promote / capture sync record the toolset it was actually
offered. Writing a suite by hand instead? Inline tools: directly on each
example — see Agent migrations. If your
tools are defined as LangChain Tool objects, export them to JSON Schema
once for that inline list. Framework-side agent timelines can also be
scored via external traces.
Yes — one suite example per turn, each carrying a recorded history
prefix that is replayed teacher-forced: both models see byte-identical
context, and only the current turn's output is compared. See
Multi-turn conversations. Full-conversation
re-driving (feeding the candidate's own replies into later turns) is
deliberately not supported — it breaks the paired-comparison contract.
Because the two defaults differ on purpose. The library default for
every evaluator's blocking is true, so a hand-written evalshift.yaml
that lists semantic: (or an llm_judge entry) without the key gates the
verdict on it. evalshift init writes blocking: false for both: the
semantic score measures drift from the source output, not correctness,
so a right answer in different words reads as a regression, and on the
small suites a fresh capture starts with that noise (and judge noise)
would dominate the verdict. Use a judge criterion for correctness.
The library default is not flipped to match because that would silently
turn a failing migration into a passing one for every existing config that
relies on the omitted key — a gate loosened under a minor release. It
stays true until a version: 2 schema. To get init's behaviour in a
hand-written file, say so:
evaluators:
semantic:
embedding_model: text-embedding-3-small
blocking: false
llm_judge:
- criterion_name: helpfulness
criterion_prompt: Which answer helps the user more?
blocking: falseSee blocking.
They're hidden — they're development aids, not part of the supported
user pipeline. Both still run if you invoke them by name. validate also
prints the CI pin drift warning after its
success line (advisory; the exit code is unchanged).