In-process capture SDK for EvalShift.
Install it inside your agent process to record what the agent does — model calls, tool calls,
retrievals — and write CLI-valid traces to .evalshift/captures/. The evalshift CLI reads
those captures from disk; the SDK and CLI never call each other.
- Distribution:
evalshift-sdk· import name:evalshift - Runtime deps: none (stdlib-only)
- Python: >= 3.10
- Capture is off by default — set
EVALSHIFT_CAPTURE=1to record. - License: MIT
Point your coding agent at the dense, single-file reference for the piece it is working on:
- EvalShift CLI: https://www.evalshift.dev/cli-llms-full.txt
- EvalShift SDK: https://www.evalshift.dev/sdk-llms-full.txt (source of truth: llms-full.txt in this repo)
- EvalShift GitHub Action (CI): https://www.evalshift.dev/ci-llms-full.txt
pip install evalshift-sdk
# or
uv add evalshift-sdkOptional LangChain integration (EvalShiftCallbackHandler):
pip install "evalshift-sdk[langchain]" # adds langchain-core>=0.2Optional provider client wrappers (wrap_openai / wrap_anthropic / wrap_genai):
pip install "evalshift-sdk[openai]" # openai>=1.40
pip install "evalshift-sdk[anthropic]" # anthropic>=0.40
pip install "evalshift-sdk[google-genai]" # google-genai>=1.0Every adapter module is import-guarded, so the SDK stays dependency-free at runtime unless you opt in.
Co-install note: the EvalShift CLI (PyPI
evalshift, import packageevalshift_cli) depends on this SDK, so both live in one environment andpip install evalshiftbrings the SDK with it. Production agents that only record captures installevalshift-sdkalone.
from evalshift import capture
@capture.agent(suite="support_agent", redact=True, tools=[]) # no-op unless EVALSHIFT_CAPTURE=1
def handle_ticket(query): ...Already calling a provider SDK directly? Wrap the client once and every call inside the agent records itself — model, tools offered, tool calls requested, usage, latency:
from openai import OpenAI
from evalshift.adapters.openai import wrap_openai # also: wrap_anthropic, wrap_genai
client = wrap_openai(OpenAI()) # OpenAI(base_url=...) covers DeepSeek, Ollama, vLLM, Groq, ...Full guide: DOCS.md · dense LLM reference: https://www.evalshift.dev/sdk-llms-full.txt · locked design decisions: docs/DECISIONS.md
Capture writes one JSON file per sampled invocation, so .evalshift/captures/ is kept bounded by
default (no configuration needed): identical-input re-runs are de-duplicated, and each suite
directory is capped at the 200 newest captures (oldest evicted). Tune it with env vars — no code
change required (precedence: an explicit configure(...) call > env var > built-in default):
| Env var | Default | Meaning |
|---|---|---|
EVALSHIFT_MAX_CAPTURES |
200 |
Max captures kept per suite dir. 0 / none / unlimited = uncapped. |
EVALSHIFT_DEDUP |
on |
Collapse identical-input captures (per-process). off to disable. |
EVALSHIFT_CAPTURE_TTL |
off | Evict captures older than N seconds. |
EVALSHIFT_SAMPLE_RATE |
off | Capture only this fraction of runs, e.g. 0.25. |
EVALSHIFT_DIR |
.evalshift |
Capture root directory. |
A malformed value falls back to the default (capture never crashes). To restore fully unbounded
capture: EVALSHIFT_MAX_CAPTURES=0 EVALSHIFT_DEDUP=off. Disable dedup with off (or 0/none);
false/no are not recognised and leave dedup on. EVALSHIFT_SAMPLE_RATE=0 means sampling off
(capture every run) — it is not the same as configure(sample_rate=0.0), which captures nothing.
The same knobs are available in code via
configure(max_captures=..., dedup=..., capture_ttl=..., sample_rate=...).
- Build a golden eval suite from production traffic — the capture → promote loop this SDK feeds.
- Evaluating agent tool calls: what text evals can't see — why the captured tool calls matter more than the output text.
- Capture SDK docs