feat(run): serve tool-calling examples from the response cache - #26
Merged
Merged
Conversation
The cache row could only hold response text, so a tool-calling response (calls with ids and arguments, refusal, round count) had nowhere to live. Add a nullable trace_json column holding ToolTrace JSON; CachedResponse and CacheStore.put gain a `trace` field that round-trips it exactly. A row whose trace no longer validates reads as a miss. The finish_reason backfill becomes a generic additive-column backfill, so a cache.db written by 1.1.0 gains the column on open and keeps serving its text rows; the 1.1.0 code reads a migrated DB unchanged. Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
run, compare, evaluate and the integration pipelines open the cache without a URL, which resolves to ~/.evalshift/cache.db. The suite read and wrote that file (test_run_command and test_compare_command left 12 rows in it), so a row from one test run could serve the next. An autouse fixture now points DEFAULT_CACHE_PATH at a per-test temp file. Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
The run stage routed every example with a toolset around the cache, so each run of an agent suite paid for every call again, one per replayed round. The bypass dated from v0.2, when a cache row could not hold a parsed tool trace; the key already had the toolset and round dimensions waiting for this. Each replayed round is now its own cache entry, keyed on the canonical model, prompt and inputs, the exact message list that round sends (history, current turn, recorded rounds and fixture results), the toolset fingerprint (strict included), the generation config (tool_choice, parallel_tool_calls), the effective temperature and max_tokens, the round index and the sample index. Teacher forcing makes rounds independent requests, so a hit feeds the unchanged merge and the raw.jsonl row is identical apart from `cached`. Policies match the text path: errored rounds are not cached, truncated rounds are cached and warned about on a hit, defaults.cache: false neither reads nor writes, and a cached round carries its original cost and latency. A row counts as cached only when every round was a hit. Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Revert the "tool-less examples only / agent suites run live at full price" statements that PR #23 added across DOCS.md, llms-full.txt and docs/{configuration,evaluators,getting-started,faq}.md, and describe the per-round key instead. docs/agents.md gains a bullet on how replayed rounds cache. CHANGELOG: a Changed entry, and the unreleased docs-audit entry now says the tool-less limit was the old behaviour. Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
complete_messages_with_tools built the provider tools array inline. The run cache is about to key on exactly that array, so it moves into a module-level serialize_tools(canonical, tools) that dispatch calls; the key and the wire then share one source. Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
The tool-round key used the toolset fingerprint, which sorts tools by name, while dispatch sends them in list order. Reordering an example's tools changed what the provider received but still hit the cache. cache_key's toolset_fingerprint (only the tool path passed it) becomes tools_payload: the serialize_tools output, hashed in order. The tool path keys on the same list the client sends. _fingerprint_toolset had no other caller and is removed. Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
… schema Opening the cache checks the schema and then changes it. Two processes opening one DB can both see it missing: the loser's CREATE fails with "table cached_calls already exists" (fresh DB, present on 1.1.0 too) or its ALTER with "duplicate column name: trace_json" (legacy DB). Either error aborts `run`. The reviewer's 4-process repro failed 70/160 and 104/160 opens; after this change both fail 0/160. Each change now runs in its own transaction and swallows exactly those two SQLite messages; any other OperationalError is raised. Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
A multi-round tool row that re-sent only some rounds is not `cached` (it spent money), but its summed latency mixes this run's rounds with the original latency of the cached ones. The live latency stats and the report and bundle latency deltas still treated it as a measurement. Call gains `cached_rounds` (defaults to 0 for older raw.jsonl rows) and a `latency_replayed` property. The orchestrator records the count on both paths. Economics' live latency stats, report.json's `latency_comparable` and the bundle's `latency_comparable` now test `latency_replayed`. The count is not uploaded, so the server contract is unchanged. The Call.latency_ms docstring said cache hits carry 0 and also that they keep the original latency. It now says what the code does: they keep the original. Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
_execute translated the example's generation_config for the cache key and _execute_with_tools translated it again for dispatch, so every "not understood" warning fired twice per tool call. _execute now passes its result through. Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
The cache docs now describe the tool-path key as the tool list exactly as sent, in order, not a toolset fingerprint. They say that a partly cached row's latency is left out of live latency figures, and that concurrent opens are safe. CHANGELOG gains a Fixed bullet for the fresh-cache open race, which 1.1.0 also has. The docs/agents.md minimal config drops a stray `cache: false` that nothing on the page explained. Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Latency means cover only calls measured fully live on this run, so a role with none of them (all cached, or every row partly cached) averages 0.0, meaning "unmeasured". The insight facts and the HTML header refused a zero source but not a zero target, so a cached target beside a live source rendered as a -100% latency change. Both now report it as not comparable. The facts comment said cache hits carry latency_ms = 0, which is wrong: they keep their original latency and are excluded via Call.latency_replayed. Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
CacheStore.open created an engine and, if the schema setup raised, left it undisposed: no store owned it yet. It now disposes the engine before re-raising. New tests cover a non-race error from create_all and one from the column backfill, and check that both propagate and dispose the engine. Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
The report's economics note now says latency covers calls measured fully live, and that cache hits and rows with any cached round are excluded. The _execute_with_tools docstring names cached_rounds among the fields a hit changes. Paragraphs this branch edited in CHANGELOG.md, llms-full.txt and two docstrings are re-wrapped. Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
babaliauskas
force-pushed
the
feat/cache-tool-calls
branch
from
October 1, 2026 12:14
c6f322a to
b55e155
Compare
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Why
runskipped the response cache for any example with a toolset. The comment said "caching serialised traces is a v0.3 polish", and the key'sround_indexslot had been reserved for this in advance. So every repeat run of an agent suite, including each round of a teacher-forced multi-round replay, was dispatched live at full price.What
One cache entry per provider round. Each round is keyed on exactly what it sends:
max_tokensgeneration_config(coverstool_choice/parallel_tool_calls)The tools array comes from a new
serialize_tools, which is also what builds the dispatched payload, so the key and the wire can't diverge. A cache hit reproduces the liveToolCompletionResult: trace, calls, tokens, cost, original latency and finish_reason, through the same round merge.Policies match the text path:
defaults.cache: falsemeans no reads and no writescachedmeans every round hitPartly cached rows. A new
Call.cached_roundsfield, not uploaded, keeps multi-round rows that were only partly cached out of live-latency stats andlatency_comparable.Cache DB:
trace_jsoncolumn, added by an additive migration; old caches keep working in both directionsAlso fixed along the way
~/.evalshift/cache.db. An autouse fixture now isolates them.generation_config"not understood" warnings no longer print twice per tool call.Docs
docs/*.mdand CHANGELOG describe the new behaviour and the key.Checks
make ciis green: 2415 tests, 94.51% coverage.🤖 Generated with Claude Code