Local trace teardown · public API rates

See what that LLM run really cost.

Bring the trace you already have. Get the re-send multiplier, the API-rate total and the fixes worth making — without uploading a word.

  • Runs in this tab
  • No account
  • No upload

Step 1

Bring one run

Use a file, paste structured data, or try the built-in example.

Your input stays in this tab and disappears on reload. How we prove it.

Why a calculator with a text box gets this wrong

A full-history API conversation is not one request. It is N requests, and request k carries every turn before it, because the API is stateless and the client re-sends the history. The transcript contains each turn once; the request stream bills the replayed history again and again. A consumer chat export can only reconstruct that stream, so its dollar figure is an API-list-price equivalent, not a subscription invoice.

Paste that conversation into a text box and it gets tokenized as a single blob — counted one time. On twenty-turn conversations built from five real corpora, the billed input was 9.7x to 11.4x the visible transcript. No amount of pasting reaches that number, because the information needed to compute it — where one call ends and the next begins — is exactly what a blob throws away.

How the curve behaves, and where it turns →

Caching is not a discount slider

Cache reads cost about a tenth of the input rate. Cache writes cost a quarter more than just sending the tokens. Every calculator we compared against either ignores the write or says it is excluded, which is what makes caching look free — and on a trace whose prefix keeps changing, caching is a straight loss.

When caching pays, and when it costs you →

Agent retries need an evidence ledger

A repeated tool call can be waste, recovery, pagination or polling. The Claude Code session audit prices identical adjacent attempts exactly, then keeps recovery work and suspected loops in separate columns so a scary total cannot outrun its evidence.

Audit a Claude Code session locally →

Where we are exact, and where we are not

OpenAI counts are exact, computed here from the real byte-pair-encoding tables. Claude and Gemini are exact too — if your trace is an API log, because the provider counted the tokens and put the number in the response. From a chat export they are estimated, and every figure says so on its face rather than in a footnote.

The full exact-versus-estimated table →