Local trace teardown · public API rates
See what that LLM run really cost.
Bring the trace you already have. Get the re-send multiplier, the API-rate total and the fixes worth making — without uploading a word.
- Runs in this tab
- No account
- No upload
Step 1
Bring one run
Use a file, paste structured data, or try the built-in example.
Drop a trace file
JSONL, JSON, CSV, TXT or LOG
02Paste a structured traceSame local engine, no file needed
See the answer first.
Run a 14-call example with no data of your own.
Which trace format gives the best answer?
- BESTAPI request log (
.jsonl) — exact for every provider when responses carry vendor usage. - AGENTClaude Code session (
.jsonl) — exact usage plus repeated-tool and recovery steps. - GOODCall log (
.csv) — exact totals, but no message-level re-send curve. - OKChatGPT or Claude export — calls reconstructed; Claude figures estimated.
- LASTLabelled transcript — reconstructed from speaker boundaries, never priced as one blob.
Your input stays in this tab and disappears on reload. How we prove it.
Reading…
Step 2 · Your teardown
The re-send is the first number to fix.
billed input in the log, or reconstructed API input for an export, against the visible transcript
- API-rate total
- —
- Calls
- —
- Basis
- —
Growth curve
What each new call carried forward
- cumulative
- this call alone
Statement
Cost at public API list prices
Ranked actions
What to change first
Cache check
Prompt caching, on this exact prefix
Agent waste map
Proven repetition, recovery and suspicion — kept separate
Method, confidence and limitsRead before budgeting
Why a calculator with a text box gets this wrong
A full-history API conversation is not one request. It is N requests, and request k carries every turn before it, because the API is stateless and the client re-sends the history. The transcript contains each turn once; the request stream bills the replayed history again and again. A consumer chat export can only reconstruct that stream, so its dollar figure is an API-list-price equivalent, not a subscription invoice.
Paste that conversation into a text box and it gets tokenized as a single blob — counted one time. On twenty-turn conversations built from five real corpora, the billed input was 9.7x to 11.4x the visible transcript. No amount of pasting reaches that number, because the information needed to compute it — where one call ends and the next begins — is exactly what a blob throws away.
How the curve behaves, and where it turns →
Caching is not a discount slider
Cache reads cost about a tenth of the input rate. Cache writes cost a quarter more than just sending the tokens. Every calculator we compared against either ignores the write or says it is excluded, which is what makes caching look free — and on a trace whose prefix keeps changing, caching is a straight loss.
When caching pays, and when it costs you →
Agent retries need an evidence ledger
A repeated tool call can be waste, recovery, pagination or polling. The Claude Code session audit prices identical adjacent attempts exactly, then keeps recovery work and suspected loops in separate columns so a scary total cannot outrun its evidence.
Audit a Claude Code session locally →
Where we are exact, and where we are not
OpenAI counts are exact, computed here from the real byte-pair-encoding tables. Claude and Gemini are exact too — if your trace is an API log, because the provider counted the tokens and put the number in the response. From a chat export they are estimated, and every figure says so on its face rather than in a footnote.