Restart-proof KV cache for llama-server: 40K-token TTFT 155s to 6s after reboot #29885
Replies: 1 comment
|
Update to the README since this went up — three things worth calling out for anyone trying it: 1. Prebuilt CUDA binaries exist for both platforms. The ./llama-server -m model.gguf -c 32768 --host 0.0.0.0 --port 8000 \
--slot-save-path /var/tmp/kvstore/model \
--prompt-cache-disk-budget 20 --checkpoint-min-step 2048 --ctx-checkpoints 642. Only one flag is actually required. 3. Cache hits require a byte-identical prompt prefix. Only the common prefix from token 0 onward is reused. If you prepend a timestamp or a random ID, everything after it misses and the cache is effectively unused. Put stable content first (system prompt, tool definitions, long docs) and per-turn content last — opencode / Claude Code already do this, so it works out of the box. The README now has a "when does the cache hit / when does it miss" section covering this, a flag-duty table, and a Chinese-first three-step guide (the repo README is now Chinese-primary). A Windows Still happy to upstream this behind a flag if maintainers are interested — the whole feature is one commit touching |
Uh oh!
There was an error while loading. Please reload this page.
Uh oh!
There was an error while loading. Please reload this page.
Every restart wipes your prefill. Server crash, box reboot, model reload — a 40K-token agentic prompt pays the full ~155s prefill again before the first token.
I patched llama-server to persist KV-cache prefix checkpoints to disk and auto-restore the longest matching prefix on boot. Same prompt after a restart: first token in ~6s instead of ~155s (2.1s disk read once + decode). No client changes, no re-architecture — the server just remembers.
Measured on GTX 1080 8GB, Qwen3.6-35B-A3B,
-c 92160:Single-commit patch (86KB): https://github.com/qqkzlm/llama-cpp-prompt-disk-cache —
git amonto the PrismMLprismbranch, four flags (--slot-save-path,--checkpoint-min-step,--ctx-checkpoints,--prompt-cache-disk-budget).One MoE lesson from the same box: keep experts on GPU, let KV follow the layers — pushing 20 MoE layers to CPU collapsed prefill 265→30 tok/s, which no cache can fix.
Happy to upstream behind a flag if maintainers are interested.
All reactions