Skip to content

Repository files navigation

llmtoll

Point your app at llmtoll instead of api.anthropic.com, and see what every call actually costs — cache hits included.

No config file. No database to stand up. One container, one environment variable, and the numbers start showing up.

CI Python 3.12+ License: MIT

English | 中文

The llmtoll dashboard

Start in 30 seconds

docker run -p 127.0.0.1:8080:8080 -v llmtoll:/data ghcr.io/skyn9/llmtoll

Then point your client at it:

export ANTHROPIC_BASE_URL=http://localhost:8080

That's the whole setup. Your existing key, your existing code — the SDK reads ANTHROPIC_BASE_URL on its own, so nothing else changes:

import anthropic

client = anthropic.Anthropic()          # picks up ANTHROPIC_BASE_URL
client.messages.create(
    model="claude-opus-5",
    max_tokens=1024,
    messages=[{"role": "user", "content": "Hello"}],
)

Open http://localhost:8080 for the dashboard, or curl localhost:8080/api/stats for JSON.

Running it without Docker
uv tool install git+https://github.com/skyn9/llmtoll
llmtoll run

llmtoll stats --window 7d prints the same figures in the terminal. Not on PyPI; install from git. Append @v0.1.0 to pin a release rather than track main.

What it records

Every call, streamed or not, lands in SQLite with:

  • All four token counters — input, output, cache writes, cache reads
  • Cost, priced per counter (see below), or explicitly unpriced if the model is unknown
  • Latency — total, plus time-to-first-token for streamed calls
  • Outcome — status, error type, stop reason, upstream request id, and whether the client hung up mid-stream
  • Who called — a label you set, or a hash of the caller's key. The key itself is never stored.

Getting the cost right

This is the part most cost trackers get wrong, so it is worth being specific.

A cached token is not an input token. Anthropic bills four counters at four different rates, all derived from the model's base input price:

Counter Rate On claude-opus-5 ($5/$25 per Mtok)
input_tokens 1× input $5.00 / Mtok
cache_creation_input_tokens 1.25× or 2× input $6.25 or $10.00 / Mtok
cache_read_input_tokens 0.1× input $0.50 / Mtok
output_tokens 1× output $25.00 / Mtok

Bill a cache read as a fresh input token and you overstate a cache-heavy workload by 10× on that portion.

Which cache-write rate applies is not in the response. The multiplier depends on whether the breakpoint was written with a 5-minute or 1-hour TTL, and the API reports only a combined cache_creation_input_tokens with no tier attached. The TTL lives in the cache_control block of the request. llmtoll reads it from there — the response alone cannot tell you.

When one request mixes both TTLs, an exact split is not derivable from what the API returns. llmtoll prices the write at the 1-hour rate and flags the record as ambiguous, so the figure is a stated upper bound rather than a confident wrong number.

An unknown model is never $0. A model missing from pricing.toml is recorded with a null cost and surfaced in the dashboard, because a pricing gap that reads as "free" is a hole you find on the invoice.

Money is integer nano-dollars, not floats. A call can cost $0.000015; summing a month of those in floating point drifts. Prices parse as Decimal, costs store as int at 1e-9 USD, and totals sum exactly.

Keeping spend under control

Once the numbers are right, two things can be done with them.

A daily ceiling.

LLMTOLL_DAILY_BUDGET_USD=50

Past the limit, calls are refused locally with a 429 and a Retry-After pointing at the next UTC midnight. Upstream never sees them, so a runaway loop costs nothing once the ceiling is hit. Refusals are recorded, so the dashboard shows what was held back rather than just going quiet.

Be clear about what this does and does not promise: it is enforced after the fact. A call's cost is not known until it finishes, so the ceiling stops the next request once the total is reached — not the one that crosses it. A $50 budget can settle at $50.04. The alternative would be reserving against max_tokens up front, which over-reserves so badly on typical traffic that it would refuse calls you can plainly afford.

The running total is restored from the database at startup, so restarting the process is not a way to get a fresh budget.

One gap worth knowing: a call llmtoll cannot price does not move the total. A model missing from pricing.toml — most likely on the day one ships — spends real money that the ceiling never sees. The dashboard says loudly when that is happening.

Several keys, rotated.

LLMTOLL_ANTHROPIC_API_KEYS=sk-ant-one,sk-ant-two,sk-ant-three

Calls go round-robin across the pool. When one comes back 429, that key leaves the rotation for as long as its Retry-After says and the request is retried on a different one — the caller sees a single successful response and never learns the first attempt happened. Transient 5xx and 529 responses are retried with exponential backoff.

Retrying is only safe in one narrow window, and the code is arranged around it: headers have arrived but the body has not been touched, so a failed attempt can be discarded whole. Once a streamed response starts reaching the client, retrying is off the table — those bytes cannot be recalled, so a mid-stream failure is reported rather than papered over.

Keys never appear in the database or the dashboard; callers are told apart by a hash, and key health shows up as …1234 on /healthz.

Why another one of these

LiteLLM, Helicone and Langfuse are all more capable than this, and if you need what they do you should use them:

llmtoll LiteLLM Helicone Langfuse
Providers Anthropic (OpenAI planned) 100+ many many
Setup one container config YAML Postgres + ClickHouse + Redis, or SaaS Postgres + ClickHouse, or SaaS
Prompt/trace inspection no some yes yes
Evals, datasets, playground no no some yes
Budgets, key rotation yes yes some no
Provider routing, load balancing no yes some no
Cache-tier-aware cost yes partial partial partial
Runs on a laptop with nothing else installed yes no no no

Use llmtoll when the question is narrow — what is this costing me, and where is it going — and you would rather not run a data platform to answer it. Move to one of the others when you need prompt-level tracing, evals, or multi-provider routing.

How it works

your app ──► llmtoll ──► api.anthropic.com
               │
               ├─ pricing engine   per-counter cost, cache-tier aware
               ├─ bounded queue    async writer, drops before it blocks
               ├─ SQLite           WAL; no server to run
               └─ dashboard        server-rendered, no build step

Three rules the code is built around:

Bytes reach the client before they reach the meter. In a streamed response each chunk is yielded downstream first and only then fed to the usage extractor. Metering adds no latency, and the client sees the same time-to-first-token it would without the proxy. There is a test that fails if this inverts — run over a real socket, because an in-process test transport buffers and would pass either way.

Telemetry can fail; the gateway cannot. Records go into a bounded in-memory queue via a call that never awaits and never raises. A background task batches them into SQLite. If the queue fills, the oldest record is dropped and counted — under load the recent past is worth more than the distant past. If the database is unreachable, the batch is dropped, the error is counted, and the writer keeps running. Nothing in this path can surface to a caller.

Partial calls still count. A client that hangs up mid-stream was still billed by upstream for whatever was generated. llmtoll records that spend and marks the call as disconnected, rather than losing it because no final response arrived.

Configuration

Every setting has a working default; all are environment variables with the LLMTOLL_ prefix.

Variable Default
LLMTOLL_PORT 8080
LLMTOLL_HOST 127.0.0.1 (0.0.0.0 in Docker)
LLMTOLL_UPSTREAM_BASE_URL https://api.anthropic.com Point at any compatible endpoint
ANTHROPIC_API_KEY unset Optional: hold the key here so client configs don't need it
LLMTOLL_ANTHROPIC_API_KEYS unset Comma-separated pool to rotate across
LLMTOLL_DAILY_BUDGET_USD unset Refuse calls past this much spend per UTC day
LLMTOLL_MAX_UPSTREAM_ATTEMPTS 3 Total tries per request, including the first
LLMTOLL_RETRY_BACKOFF_SECONDS 0.5 Doubles each retry
LLMTOLL_DATA_DIR ~/.llmtoll (/data in Docker)
LLMTOLL_PRICING_PATH bundled pricing.toml Point at your own to override rates
LLMTOLL_QUEUE_MAXSIZE 10000 Records buffered before the oldest is dropped
LLMTOLL_UPSTREAM_TIMEOUT 900 Seconds; a max-effort call legitimately runs for minutes

Send x-llmtoll-key: <label> on a request to name the caller in the dashboard. Without it, callers are distinguished by a hash of their API key. The label is whatever the caller says it is — it attributes cost, it does not authenticate.

Who can reach it

llmtoll has no authentication. That is fine for the thing it is — a meter you run next to your own work — and it is why the default bind address is 127.0.0.1. It stops being fine the moment the port is reachable by anyone else, because whoever reaches it can:

  • spend the key you configured, without holding a key of their own;
  • read the dashboard: your models, your callers, your spend;
  • see the last four characters of each pooled key on /healthz.

The container binds 0.0.0.0 because nothing else works inside a container, so the confinement has to happen at the published port — hence -p 127.0.0.1:8080:8080 above. Publishing it more widely than that means putting something in front of it that authenticates.

Keeping prices current

Rates live in pricing.toml, separate from the code, with promotional windows and fast-mode tiers as first-class fields. When a model shows up in traffic that the table does not cover, the dashboard says so instead of quietly booking it at zero.

Development

uv sync
uv run pytest                                  # 120 tests
uv run ruff check src tests scripts
uv run mypy                                    # strict
uv run python scripts/seed_demo.py             # synthetic data for screenshots

scripts/seed_demo.py writes made-up records to a scratch database. It calls no API and spends nothing — point it somewhere disposable.

Status

Early, but the Anthropic path is complete and tested: streaming and non-streaming proxying, cache-aware costing, daily budgets, and key rotation with 429/529 failover. Next up: an OpenAI-compatible endpoint, Prometheus /metrics, and response caching. Issues and PRs welcome.

License

MIT

About

A zero-config self-hosted gateway that meters every LLM call: tokens, cache-aware cost, latency, errors.

Topics

Resources

Stars

1 star

Watchers

0 watching

Forks

Releases

Packages

Contributors

Languages