A full-history 20-turn API conversation bills about 10x its visible text

The API keeps nothing between calls. When a client sends the whole conversation again every turn, here is what that does to billed input.

Why it happens

A chat completion endpoint is stateless. It does not remember your previous message, and it has no session to look up. So the only way the model can know what was said three turns ago is for your client to send it again — along with everything else, on every request.

For a full-history client, call 1 sends one turn, call 2 sends three, call 3 sends five, and call k sends 2k−1. The transcript grows one turn at a time; billed input grows by the whole conversation each time. Add the turns up and total input is roughly proportional to the square of the turn count, while the text you can read grows in a straight line.

What that costs, measured

Twenty-turn conversations were built from five real corpora — Markdown prose, contract prose, JavaScript, Python and JSON — and priced with the real tokenizer:

Billed input against visible transcript, 20-turn conversations, measured 2026-08-13.
ContentFull-history billed input
Markdown prose11.4x
JavaScript10.6x
JSON10.1x
Contract prose9.8x
Python9.7x

Ten times. Not ten percent — ten times the transcript under the stated full-history assumption, for a conversation of an ordinary length. A real request log proves whether the client actually behaved that way; an export only reconstructs it.

Where the curve turns

There is a specific call on every conversation where you stop paying for the conversation and start paying to remind the model of it: the first call at which the re-sent history costs more than that call's own new content.

On most real traces it arrives embarrassingly early — usually turn two or three. Everything after it is money spent on repetition. The teardown marks that exact call on your own chart.

When it is NOT quadratic, which matters

If your prompt is dominated by something large and fixed — a big system prompt, a retrieved document set, a long tool schema — then the constant swamps the growth and your cost is close to linear. That is a completely different problem with a completely different fix: shrink the constant, not the history.

So this site does not tell you your costs are quadratic. It measures the growth exponent of your trace and reports the number. Around 1.0 means a fixed prompt dominates. Approaching 2.0 means the history does. Most real traces land in between, and which end you are on decides which lever is worth pulling.

What to do about it

  • Truncate the history. Keep the last K turns instead of all of them. This is the direct fix, and it costs you the model's ability to refer to anything that fell off the window.
  • Summarise instead of truncating. Replace old turns with a compact summary. Cheaper than the full history, more faithful than dropping it.
  • Shrink the system prompt. It rides on every single call, so every token you cut is multiplied by the call count. On a 40-call run, 200 tokens removed is 8,000 tokens never sent.
  • Cache the stable prefix — but read what caching actually costs first, because on a prefix that keeps changing it loses money.

Each of these is worth a different amount on a different trace, which is the whole reason to measure yours instead of reading a list.

Run this on your own trace →