Skip to content

feat(run): serve tool-calling examples from the response cache - #26

Merged
babaliauskas merged 14 commits into
mainfrom
feat/cache-tool-calls
Oct 1, 2026
Merged

babaliauskas merged 14 commits into
mainfrom
feat/cache-tool-calls

Conversation

@babaliauskas

Copy link
Copy Markdown
Collaborator

Why

run skipped the response cache for any example with a toolset. The comment said "caching serialised traces is a v0.3 polish", and the key's round_index slot had been reserved for this in advance. So every repeat run of an agent suite, including each round of a teacher-forced multi-round replay, was dispatched live at full price.

What

One cache entry per provider round. Each round is keyed on exactly what it sends:

  • canonical model id, prompt and inputs
  • effective temperature and max_tokens
  • the exact message list for that round
  • generation_config (covers tool_choice / parallel_tool_calls)
  • the tools array exactly as sent, in order
  • round and sample index

The tools array comes from a new serialize_tools, which is also what builds the dispatched payload, so the key and the wire can't diverge. A cache hit reproduces the live ToolCompletionResult: trace, calls, tokens, cost, original latency and finish_reason, through the same round merge.

Policies match the text path:

  • errors are not cached; earlier successful rounds keep their entries
  • truncated rounds are cached and warn on hit
  • defaults.cache: false means no reads and no writes
  • cached means every round hit

Partly cached rows. A new Call.cached_rounds field, not uploaded, keeps multi-round rows that were only partly cached out of live-latency stats and latency_comparable.

Cache DB:

  • a new nullable trace_json column, added by an additive migration; old caches keep working in both directions
  • concurrent opens racing to create or migrate the schema no longer crash; this was a pre-existing bug for fresh DBs
  • the engine is disposed if schema setup fails

Also fixed along the way

  • Tests wrote to the developer's real ~/.evalshift/cache.db. An autouse fixture now isolates them.
  • "−100%" latency change. The report and the insights showed it when the target was served entirely from cache. They now say "not comparable" unless both roles measured latency live. This also affects 1.1.0.
  • Duplicate warnings. generation_config "not understood" warnings no longer print twice per tool call.

Docs

  • DOCS.md, llms-full.txt, docs/*.md and CHANGELOG describe the new behaviour and the key.
  • The "tool-calling examples are not cached" statements from docs: correct stale facts found by a docs-currency audit #23 are reverted.
  • The server spec's latency wording is fixed in evalshift-server#13.
  • The website draft follows.

Checks

  • make ci is green: 2415 tests, 94.51% coverage.
  • Independent review on opus, then two fix rounds, each re-reviewed.
  • The reviewer reproduced the migration race (86/160 opens failed). After the fix it's 0/160.

🤖 Generated with Claude Code

babaliauskas and others added 14 commits October 1, 2026 14:11
The cache row could only hold response text, so a tool-calling response
(calls with ids and arguments, refusal, round count) had nowhere to live.
Add a nullable trace_json column holding ToolTrace JSON; CachedResponse
and CacheStore.put gain a `trace` field that round-trips it exactly. A
row whose trace no longer validates reads as a miss.

The finish_reason backfill becomes a generic additive-column backfill,
so a cache.db written by 1.1.0 gains the column on open and keeps
serving its text rows; the 1.1.0 code reads a migrated DB unchanged.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
run, compare, evaluate and the integration pipelines open the cache
without a URL, which resolves to ~/.evalshift/cache.db. The suite read
and wrote that file (test_run_command and test_compare_command left 12
rows in it), so a row from one test run could serve the next. An
autouse fixture now points DEFAULT_CACHE_PATH at a per-test temp file.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
The run stage routed every example with a toolset around the cache, so
each run of an agent suite paid for every call again, one per replayed
round. The bypass dated from v0.2, when a cache row could not hold a
parsed tool trace; the key already had the toolset and round
dimensions waiting for this.

Each replayed round is now its own cache entry, keyed on the canonical
model, prompt and inputs, the exact message list that round sends
(history, current turn, recorded rounds and fixture results), the
toolset fingerprint (strict included), the generation config
(tool_choice, parallel_tool_calls), the effective temperature and
max_tokens, the round index and the sample index. Teacher forcing makes
rounds independent requests, so a hit feeds the unchanged merge and the
raw.jsonl row is identical apart from `cached`.

Policies match the text path: errored rounds are not cached, truncated
rounds are cached and warned about on a hit, defaults.cache: false
neither reads nor writes, and a cached round carries its original cost
and latency. A row counts as cached only when every round was a hit.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Revert the "tool-less examples only / agent suites run live at full
price" statements that PR #23 added across DOCS.md, llms-full.txt and
docs/{configuration,evaluators,getting-started,faq}.md, and describe the
per-round key instead. docs/agents.md gains a bullet on how replayed
rounds cache. CHANGELOG: a Changed entry, and the unreleased docs-audit
entry now says the tool-less limit was the old behaviour.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
complete_messages_with_tools built the provider tools array inline. The
run cache is about to key on exactly that array, so it moves into a
module-level serialize_tools(canonical, tools) that dispatch calls; the
key and the wire then share one source.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
The tool-round key used the toolset fingerprint, which sorts tools by
name, while dispatch sends them in list order. Reordering an example's
tools changed what the provider received but still hit the cache.

cache_key's toolset_fingerprint (only the tool path passed it) becomes
tools_payload: the serialize_tools output, hashed in order. The tool
path keys on the same list the client sends. _fingerprint_toolset had
no other caller and is removed.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
… schema

Opening the cache checks the schema and then changes it. Two processes
opening one DB can both see it missing: the loser's CREATE fails with
"table cached_calls already exists" (fresh DB, present on 1.1.0 too) or
its ALTER with "duplicate column name: trace_json" (legacy DB). Either
error aborts `run`. The reviewer's 4-process repro failed 70/160 and
104/160 opens; after this change both fail 0/160.

Each change now runs in its own transaction and swallows exactly those
two SQLite messages; any other OperationalError is raised.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
A multi-round tool row that re-sent only some rounds is not `cached`
(it spent money), but its summed latency mixes this run's rounds with
the original latency of the cached ones. The live latency stats and
the report and bundle latency deltas still treated it as a
measurement.

Call gains `cached_rounds` (defaults to 0 for older raw.jsonl rows)
and a `latency_replayed` property. The orchestrator records the count
on both paths. Economics' live latency stats, report.json's
`latency_comparable` and the bundle's `latency_comparable` now test
`latency_replayed`. The count is not uploaded, so the server contract
is unchanged.

The Call.latency_ms docstring said cache hits carry 0 and also that
they keep the original latency. It now says what the code does: they
keep the original.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
_execute translated the example's generation_config for the cache key
and _execute_with_tools translated it again for dispatch, so every
"not understood" warning fired twice per tool call. _execute now passes
its result through.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
The cache docs now describe the tool-path key as the tool list exactly
as sent, in order, not a toolset fingerprint. They say that a partly
cached row's latency is left out of live latency figures, and that
concurrent opens are safe. CHANGELOG gains a Fixed bullet for the
fresh-cache open race, which 1.1.0 also has. The docs/agents.md minimal
config drops a stray `cache: false` that nothing on the page explained.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Latency means cover only calls measured fully live on this run, so a
role with none of them (all cached, or every row partly cached)
averages 0.0, meaning "unmeasured". The insight facts and the HTML
header refused a zero source but not a zero target, so a cached target
beside a live source rendered as a -100% latency change. Both now
report it as not comparable. The facts comment said cache hits carry
latency_ms = 0, which is wrong: they keep their original latency and
are excluded via Call.latency_replayed.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
CacheStore.open created an engine and, if the schema setup raised,
left it undisposed: no store owned it yet. It now disposes the engine
before re-raising. New tests cover a non-race error from create_all and
one from the column backfill, and check that both propagate and dispose
the engine.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
The report's economics note now says latency covers calls measured
fully live, and that cache hits and rows with any cached round are
excluded. The _execute_with_tools docstring names cached_rounds among
the fields a hit changes. Paragraphs this branch edited in CHANGELOG.md,
llms-full.txt and two docstrings are re-wrapped.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
@babaliauskas
babaliauskas force-pushed the feat/cache-tool-calls branch from c6f322a to b55e155 Compare October 1, 2026 12:14
@babaliauskas
babaliauskas merged commit 90eb761 into main Oct 1, 2026
4 checks passed
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant