EvalShift can compare externally recorded agent timelines. The traces come from
your own runtime — EvalShift never runs your agent — and are imported into a
completed local run, after which evaluate, analyze, and report work as
usual. Imported traces stay local: evaluate scores them and the local report
shows them, but push uploads only the replay's own tool-call trace (see
Hosted).
Run the normal model calls first so EvalShift has a run id and source/target call pairs:
evalshift run --yes
evalshift traces import <run-id> \
--source source-traces.jsonl \
--target target-traces.jsonl
evalshift evaluate <run-id>
evalshift analyze <run-id>
evalshift report <run-id> --open--source and --target are both required. Add --strict to fail the import
when any completed run pair lacks a trace pair; without it those pairs are only
counted in the printed missing pairs line.
evalshift traces import <run-id> \
--source source-traces.jsonl \
--target target-traces.jsonl \
--strictThe import command writes:
.evalshift/runs/<run-id>/traces.jsonl
Each line is one trace for one (prompt_id, example_id, role):
{"run_id":"r_20260609_trace1","prompt_id":"support_agent","example_id":"refund_017","role":"source","events":[{"type":"tool_call","sequence_index":0,"timestamp":"2026-06-09T12:00:00Z","metadata":{},"name":"check_refund_policy","arguments":{}},{"type":"tool_call","sequence_index":1,"timestamp":"2026-06-09T12:00:01Z","metadata":{},"name":"issue_refund","arguments":{"ticket_id":"T-1032"}}]}The line itself carries run_id, prompt_id, example_id (all non-empty),
role (source or target), and events (defaults to []).
Every event has type, sequence_index (integer, ≥ 0), timestamp, and
metadata (defaults to {}). Per type, on top of those:
type |
Required | Optional (default) |
|---|---|---|
model_call |
model_id (non-empty) |
input, output (null), input_tokens, output_tokens, latency_ms (0), cost_usd (0.0), toolset_ref, tools_offered, requested_tool_calls (null) |
tool_call |
name (non-empty) |
arguments ({}), call_id, parent_call_id (null) |
tool_result |
name (non-empty) |
call_id, result, error (null) |
retrieval |
source (non-empty) |
query (""), documents ([]) |
guardrail |
name (non-empty), verdict (pass / fail / warn / skipped) |
reason (null) |
final_output |
— | text ("") |
error |
message (non-empty) |
category (null) |
timestamp must carry a UTC offset — 2026-06-09T12:00:00Z or
2026-06-09T14:00:00+02:00. An offset timestamp is converted to UTC on import;
a naive one (no offset at all) is rejected with the file and line that carried
it. Assuming UTC for a naive value would silently relabel a trace recorded in
another zone, and the run bundle would then carry that as fact — the hosted
bundle contract requires UTC, so the ambiguity has to be resolved by the person
who has the trace, not by the CLI.
Numeric fields cannot be negative. Trace models are strict, like the rest of the config contract: an unknown key anywhere in the line is an error, not a warning.
A model_call distinguishes three things that are easy to conflate:
- offered —
toolset_ref(content-addressed pointer to the toolset sidecar) andtools_offered(the display-only tool-name list): what was passed to the model; - requested —
requested_tool_calls, a list of{"name": ..., "arguments": {...}, "call_id": ...}(argumentsdefaults to{},call_idtonull): what the model asked to call in its response; - executed — the
tool_call/tool_resultevents: what the application actually ran.
Requested and executed legitimately differ — an app may filter, re-order, or
wrap the calls it runs. requested_tool_calls is null on a trace recorded
before the field existed, which is not the same as [] ("the model requested
no tools").
EvalShift sorts events by sequence_index and rejects duplicate indices. A
tool_result with a call_id must match a tool_call that carried the same
call_id earlier in the trace.
evaluators:
agent_trace:
- name: trace_safety
check_tool_order: true
check_arguments: true
check_missing_verification: true
verification_tools: ["check_refund_policy"]
dangerous_tools: ["issue_refund"]The evaluator emits normal scores.jsonl records with failure categories such
as TOOL_ORDER_DRIFT, ARGUMENT_VALUE_DRIFT, DANGEROUS_ACTION_DRIFT, and
MISSING_VERIFICATION_STEP.
agent_trace compares the target with the source: dangerous_tools flags a
target that made more dangerous calls than the source did, so a dangerous call
both sides made goes unflagged. To hold both sides to rules instead, add a
trace_invariants entry with
traces: imported:
evaluators:
trace_invariants:
- name: refund_contract
traces: imported
rules:
- id: policy-before-refund
type: order
before: check_refund_policy
after: issue_refund
- id: one-refund
type: call_count
tool: issue_refund
max_calls: 1An imported trace is checked as one response with no prior context: every call
in it, in sequence_index order, was the agent's own. A rule the target breaks
counts toward migration_policy.max_invariant_violations whatever the source
did.
Debug commands become trace-aware:
evalshift diff case <run-id> <example-id>
evalshift inspect case <run-id> <example-id>
evalshift replay case <run-id> <example-id> --model target --traceImported traces are not uploaded. They stay in
.evalshift/runs/<run-id>/traces.jsonl, where evaluate (the agent_trace
evaluator and trace_invariants entries with traces: imported), the local
report and diff case / inspect case / replay case read them.
evalshift bundle / evalshift push carry only the replay's own tool-call
trace — one stream per model side with the tool calls (names, arguments, call
ids), round markers, any final text and refusal messages, capped at 256 KB per
side. The full bundle contract is in
Hosted EvalShift.