Prompt caching is not free, and on a changing prefix it loses money
Reads are cheap. Writes cost more than not caching at all. Whether you come out ahead depends entirely on how stable your prefix is — which only your own trace knows.
Cache-miss forensics
Find the first prefix break
Use a structured request log for model, system, tools and message evidence.
Drop a trace
JSONL request log or Claude Code session
The comparison happens here. Prefix text is never shown, uploaded or copied.
Comparing request prefixes…
Trace diagnosis
The first break, not generic cache advice
- First structural divergence
- —
- Missed-cache opportunity
- —
- Break-even
- —
- TTL evidence
- —
Model, system, tool order or message prefix.
Rewrite rate minus cache-read rate on proved lost prefix tokens.
Successful reads needed to repay one write premium.
Expiry is separate from structural divergence; eviction remains unknown.
The three rates, not one
Almost every cost calculator models caching as a single discount: "assume 40% of your input is cached, bill it at 10%". That is one of three numbers, and it is the flattering one.
| Model | Send it | Read from cache | Write to cache |
|---|---|---|---|
| GPT-5.6 Sol | $5.00 | $0.50 | $6.25 |
| GPT-5.6 Terra | $2.00 | $0.20 | $2.50 |
| Claude Opus 5 | $5.00 | $0.50 | $6.25 |
| Claude Sonnet 5 | $2.00 | $0.20 | $2.50 |
| Claude Haiku 4.5 | $1.00 | $0.10 | $1.25 |
Sources: OpenAI pricing and Anthropic pricing.
The pattern is consistent at both vendors: a cache read costs a tenth of sending the tokens, and a cache write costs a quarter more than sending them. Caching is a bet — you pay a premium up front for the chance to pay a tenth later.
When the bet loses
You win if the same prefix is read many times. You lose if it is written and then never read, and you lose badly if that happens on every call.
Traces where caching costs more than it saves, all of them common:
- An agent that injects fresh tool output at the top of the prompt. The prefix changes every turn, so every turn writes a new one at 1.25x and reads almost nothing.
- A retrieval prompt with rotating context. Different documents each call means a different prefix each call.
- Anything with a timestamp near the start of the system prompt. One changing token at position 40 invalidates everything after it.
- Short conversations. Two or three calls do not amortise the write premium.
- Calls spaced further apart than the cache lives. Anthropic's published rates are for a five-minute window. Two calls eleven minutes apart share no cache however identical they look.
The teardown checks the timestamps in your trace against the published window and counts the expiries, because a cache that has already gone is a write you paid for and never got back.
Google bills cached context as rent
Gemini adds a mechanic neither of the others has: on top of the read discount, cached context is charged as storage, per million tokens per hour it sits in the cache — $0.50 per MTok-hour, rising to $1.00 on 1 January 2027.
An agent that caches a 100,000-token prefix and runs for eight hours is paying rent on it the whole time, whether or not it reads the cache again. No slider models that, and this page will not pretend to price it for you either: only you know how long your context actually sat there.
What we refuse to guess
Two models in our table publish a cache-read rate but no cache-write rate — GPT-5.3 Codex and Gemini 3.7 Flash. The write is 1.25x input at every vendor that does publish it, so assuming 1.25x here would look reasonable.
We do not. The write rate is the number that decides whether caching pays, so estimating it would mean deciding the answer by assumption and then presenting it as a measurement. For those models the lever is withheld, with the reason shown in its place. An empty row that says "cannot be computed" is worth more than a confident row that is invented.
How to find out which case you are in
Drop your trace in. The teardown walks your calls in order, works out how much of each prompt is a prefix the previous call already established, prices the reads and the writes separately at the published rates, checks the gaps against the TTL, and tells you whether caching would have saved you money or cost you money on that exact run.