What this is exact about, and what it only estimates

A confidently wrong cost figure would end this site's usefulness on the first check somebody ran. So every number says which kind it is.

Token counts, by provider

ProviderFrom an API logFrom a chat export or transcript
OpenAIexact counted locallyexact counted locally
Anthropicexact the vendor's own countestimated
Googleexact the vendor's own countestimated

Why OpenAI is exact

We ship the real byte-pair-encoding rank tables — o200k_base and cl100k_base, the same data OpenAI's own tiktoken uses — and run them in your browser. That is an exact count, offline, with no request to anybody.

Why an API log is exact for everyone

A log of real requests contains the responses, and every provider puts its own token counts in the response: prompt_tokens and completion_tokens at OpenAI, input_tokens and cache_read_input_tokens at Anthropic, promptTokenCount at Google. The vendor counted the tokens and wrote the number down. We just read it.

This is why the drop zone ranks an API log above a chat export instead of treating them as two equally good ways in. It is the difference between a number that is right and a number that is close.

Why Claude and Gemini are estimated from raw text

Because there is no current local tokenizer for either. Anthropic's published package is at version 0.0.4, was last released in July 2023, and encodes Claude 1 and 2 — shipping it would put a three-year-old vocabulary behind a current model's dollar figure. Google ships no local tokenizer at all. Both vendors count server-side, and calling their endpoints would mean sending them your trace, which is the one thing this site exists not to do.

So the text is counted with o200k_base as a stated proxy, and every figure derived from it is labelled estimated. Treat those dollar amounts as an order of magnitude, not a bill.

What we measured about the estimate, and what we corrected

Twenty-turn conversations were built from five real corpora and priced twice, once with each tokenizer we hold:

FigureMoved by, across two tokenizers
Total billed input0.026% – 0.755%
The billed-to-visible multiplier0.038% – 0.182%

The multiplier's worst case is about four times tighter than the absolute figure's. An earlier version of this page claimed seventeen times. That number came from a script that computed the visible transcript by a different route than the shipped code does, so the two halves of the ratio moved together and flattered it; the site's own test contradicted it within the hour and also exposed a real double-counting bug. The figure above is the corrected one — smaller, less impressive, and true.

The limit of that measurement

Both tokenizers are OpenAI's, trained on related data. Their agreement bounds how far two related vocabularies drift. It is not evidence about how far Claude's or Gemini's vocabulary sits from OpenAI's, and we have no instrument that would tell us without sending your trace to the vendor.

So this site publishes no cross-vendor error band. Inventing one would be the same failure as a confidently wrong price, wearing a lab coat.

What survives the admission

The structural findings — the multiplier, the turning point, the input/output split, the levers — are ratios of two counts taken with the same tokenizer. The tokenizer largely cancels, and what is left is a property of your conversation: that call 20 carries twenty turns of history is true however you count them.

So "the full-history reconstruction processes 11x the visible text" remains structurally stable even for a vendor we cannot tokenize. "That was $4.12 on your invoice" does not. The result page separates them on exactly this line.

Where the prices come from

One data file, hand-maintained, small on purpose. It holds the models a real production trace actually names, and every row carries the URL it was read from and the date it was read.

Price table last verified 2026-08-13. The result page computes and shows how many days ago that was, every time it renders a figure.

Why we do not reuse the maintained table on npm

We tried to. The obvious dependency ships a price field for 109 models, and reading the values rather than the field names ended that plan: three different current models carried one identical price that belonged to none of them, five carried no price at all, and exactly one of the 109 carried a cache-read rate. Since what caching costs is half of what this site reports, a table that can price one model's cache read is not a table this site can use.

A missing rate here renders as missing. It never quietly becomes zero, and a lever that cannot be priced honestly is withheld with its reason instead of estimated.

What this site is not

  • Not a token counter. Several free ones exist and they are good at it. This one needs the whole run, not one message.
  • Not a live billing feed. It reads a file you already have. It has no API keys and no access to your account.
  • Not a recommendation. A cheaper model that needs two attempts is not cheaper. The levers show price differences on your numbers; whether the output is still good enough is yours to judge.

Back to the teardown →