Point your app at llmtoll instead of api.anthropic.com, and see what every call actually costs — cache hits included.
No config file. No database to stand up. One container, one environment variable, and the numbers start showing up.
English | 中文
docker run -p 127.0.0.1:8080:8080 -v llmtoll:/data ghcr.io/skyn9/llmtollThen point your client at it:
export ANTHROPIC_BASE_URL=http://localhost:8080That's the whole setup. Your existing key, your existing code — the SDK reads
ANTHROPIC_BASE_URL on its own, so nothing else changes:
import anthropic
client = anthropic.Anthropic() # picks up ANTHROPIC_BASE_URL
client.messages.create(
model="claude-opus-5",
max_tokens=1024,
messages=[{"role": "user", "content": "Hello"}],
)Open http://localhost:8080 for the dashboard, or curl localhost:8080/api/stats for JSON.
Running it without Docker
uv tool install git+https://github.com/skyn9/llmtoll
llmtoll runllmtoll stats --window 7d prints the same figures in the terminal.
Not on PyPI; install from git. Append @v0.1.0 to pin a release rather than
track main.
Every call, streamed or not, lands in SQLite with:
- All four token counters — input, output, cache writes, cache reads
- Cost, priced per counter (see below), or explicitly unpriced if the model is unknown
- Latency — total, plus time-to-first-token for streamed calls
- Outcome — status, error type, stop reason, upstream request id, and whether the client hung up mid-stream
- Who called — a label you set, or a hash of the caller's key. The key itself is never stored.
This is the part most cost trackers get wrong, so it is worth being specific.
A cached token is not an input token. Anthropic bills four counters at four different rates, all derived from the model's base input price:
| Counter | Rate | On claude-opus-5 ($5/$25 per Mtok) |
|---|---|---|
input_tokens |
1× input | $5.00 / Mtok |
cache_creation_input_tokens |
1.25× or 2× input | $6.25 or $10.00 / Mtok |
cache_read_input_tokens |
0.1× input | $0.50 / Mtok |
output_tokens |
1× output | $25.00 / Mtok |
Bill a cache read as a fresh input token and you overstate a cache-heavy workload by 10× on that portion.
Which cache-write rate applies is not in the response. The multiplier depends
on whether the breakpoint was written with a 5-minute or 1-hour TTL, and the API
reports only a combined cache_creation_input_tokens with no tier attached. The
TTL lives in the cache_control block of the request. llmtoll reads it from
there — the response alone cannot tell you.
When one request mixes both TTLs, an exact split is not derivable from what the API returns. llmtoll prices the write at the 1-hour rate and flags the record as ambiguous, so the figure is a stated upper bound rather than a confident wrong number.
An unknown model is never $0. A model missing from pricing.toml is recorded
with a null cost and surfaced in the dashboard, because a pricing gap that reads
as "free" is a hole you find on the invoice.
Money is integer nano-dollars, not floats. A call can cost $0.000015; summing
a month of those in floating point drifts. Prices parse as Decimal, costs store
as int at 1e-9 USD, and totals sum exactly.
Once the numbers are right, two things can be done with them.
A daily ceiling.
LLMTOLL_DAILY_BUDGET_USD=50Past the limit, calls are refused locally with a 429 and a Retry-After pointing
at the next UTC midnight. Upstream never sees them, so a runaway loop costs nothing
once the ceiling is hit. Refusals are recorded, so the dashboard shows what was
held back rather than just going quiet.
Be clear about what this does and does not promise: it is enforced after the
fact. A call's cost is not known until it finishes, so the ceiling stops the
next request once the total is reached — not the one that crosses it. A $50
budget can settle at $50.04. The alternative would be reserving against
max_tokens up front, which over-reserves so badly on typical traffic that it
would refuse calls you can plainly afford.
The running total is restored from the database at startup, so restarting the process is not a way to get a fresh budget.
One gap worth knowing: a call llmtoll cannot price does not move the total. A model
missing from pricing.toml — most likely on the day one ships — spends real money
that the ceiling never sees. The dashboard says loudly when that is happening.
Several keys, rotated.
LLMTOLL_ANTHROPIC_API_KEYS=sk-ant-one,sk-ant-two,sk-ant-threeCalls go round-robin across the pool. When one comes back 429, that key leaves
the rotation for as long as its Retry-After says and the request is retried on a
different one — the caller sees a single successful response and never learns the
first attempt happened. Transient 5xx and 529 responses are retried with
exponential backoff.
Retrying is only safe in one narrow window, and the code is arranged around it: headers have arrived but the body has not been touched, so a failed attempt can be discarded whole. Once a streamed response starts reaching the client, retrying is off the table — those bytes cannot be recalled, so a mid-stream failure is reported rather than papered over.
Keys never appear in the database or the dashboard; callers are told apart by a
hash, and key health shows up as …1234 on /healthz.
LiteLLM, Helicone and Langfuse are all more capable than this, and if you need what they do you should use them:
| llmtoll | LiteLLM | Helicone | Langfuse | |
|---|---|---|---|---|
| Providers | Anthropic (OpenAI planned) | 100+ | many | many |
| Setup | one container | config YAML | Postgres + ClickHouse + Redis, or SaaS | Postgres + ClickHouse, or SaaS |
| Prompt/trace inspection | no | some | yes | yes |
| Evals, datasets, playground | no | no | some | yes |
| Budgets, key rotation | yes | yes | some | no |
| Provider routing, load balancing | no | yes | some | no |
| Cache-tier-aware cost | yes | partial | partial | partial |
| Runs on a laptop with nothing else installed | yes | no | no | no |
Use llmtoll when the question is narrow — what is this costing me, and where is it going — and you would rather not run a data platform to answer it. Move to one of the others when you need prompt-level tracing, evals, or multi-provider routing.
your app ──► llmtoll ──► api.anthropic.com
│
├─ pricing engine per-counter cost, cache-tier aware
├─ bounded queue async writer, drops before it blocks
├─ SQLite WAL; no server to run
└─ dashboard server-rendered, no build step
Three rules the code is built around:
Bytes reach the client before they reach the meter. In a streamed response each chunk is yielded downstream first and only then fed to the usage extractor. Metering adds no latency, and the client sees the same time-to-first-token it would without the proxy. There is a test that fails if this inverts — run over a real socket, because an in-process test transport buffers and would pass either way.
Telemetry can fail; the gateway cannot. Records go into a bounded in-memory queue via a call that never awaits and never raises. A background task batches them into SQLite. If the queue fills, the oldest record is dropped and counted — under load the recent past is worth more than the distant past. If the database is unreachable, the batch is dropped, the error is counted, and the writer keeps running. Nothing in this path can surface to a caller.
Partial calls still count. A client that hangs up mid-stream was still billed by upstream for whatever was generated. llmtoll records that spend and marks the call as disconnected, rather than losing it because no final response arrived.
Every setting has a working default; all are environment variables with the
LLMTOLL_ prefix.
| Variable | Default | |
|---|---|---|
LLMTOLL_PORT |
8080 |
|
LLMTOLL_HOST |
127.0.0.1 (0.0.0.0 in Docker) |
|
LLMTOLL_UPSTREAM_BASE_URL |
https://api.anthropic.com |
Point at any compatible endpoint |
ANTHROPIC_API_KEY |
unset | Optional: hold the key here so client configs don't need it |
LLMTOLL_ANTHROPIC_API_KEYS |
unset | Comma-separated pool to rotate across |
LLMTOLL_DAILY_BUDGET_USD |
unset | Refuse calls past this much spend per UTC day |
LLMTOLL_MAX_UPSTREAM_ATTEMPTS |
3 |
Total tries per request, including the first |
LLMTOLL_RETRY_BACKOFF_SECONDS |
0.5 |
Doubles each retry |
LLMTOLL_DATA_DIR |
~/.llmtoll (/data in Docker) |
|
LLMTOLL_PRICING_PATH |
bundled pricing.toml |
Point at your own to override rates |
LLMTOLL_QUEUE_MAXSIZE |
10000 |
Records buffered before the oldest is dropped |
LLMTOLL_UPSTREAM_TIMEOUT |
900 |
Seconds; a max-effort call legitimately runs for minutes |
Send x-llmtoll-key: <label> on a request to name the caller in the dashboard.
Without it, callers are distinguished by a hash of their API key. The label is
whatever the caller says it is — it attributes cost, it does not authenticate.
llmtoll has no authentication. That is fine for the thing it is — a meter you
run next to your own work — and it is why the default bind address is 127.0.0.1.
It stops being fine the moment the port is reachable by anyone else, because
whoever reaches it can:
- spend the key you configured, without holding a key of their own;
- read the dashboard: your models, your callers, your spend;
- see the last four characters of each pooled key on
/healthz.
The container binds 0.0.0.0 because nothing else works inside a container, so the
confinement has to happen at the published port — hence -p 127.0.0.1:8080:8080
above. Publishing it more widely than that means putting something in front of it
that authenticates.
Rates live in pricing.toml, separate from the code,
with promotional windows and fast-mode tiers as first-class fields. When a model
shows up in traffic that the table does not cover, the dashboard says so instead
of quietly booking it at zero.
uv sync
uv run pytest # 120 tests
uv run ruff check src tests scripts
uv run mypy # strict
uv run python scripts/seed_demo.py # synthetic data for screenshotsscripts/seed_demo.py writes made-up records to a scratch database. It calls no
API and spends nothing — point it somewhere disposable.
Early, but the Anthropic path is complete and tested: streaming and non-streaming
proxying, cache-aware costing, daily budgets, and key rotation with 429/529
failover. Next up: an OpenAI-compatible endpoint, Prometheus /metrics, and
response caching. Issues and PRs welcome.
MIT