Claude thinking tokens: what they are and how they are billed
Thinking tokens are billed as output, they count against max_tokens, and on a 4,000 in / 900 out prompt they can account for more than 90% of the bill at high effort levels.
Thinking tokens are output tokens. The model generates them while it works through a problem, they bill at the output rate — the expensive column — and they count against max_tokens alongside the answer text. There is no separate, cheaper bucket for internal work.
That single fact explains most of the billing surprises people hit when they raise the effort level. On a 4,000 in / 900 out prompt running on Claude Sonnet 5.5, moving from low to max takes the cost from about $0.017 to about $0.413 per call. At max, roughly 96% of that is thinking you may never see on screen.
What a thinking token actually is
Thinking tokens are the tokens a model emits while reasoning before it commits to an answer. They are produced the same way answer tokens are: one at a time, by the same model, through the same decoder.
They are not a separate resource. There is no reasoning credit and no internal-token discount. Whatever the model writes, in whatever role, is output, and output has one price.
Effort is the dial that controls how many of them get produced. At low the model does very little visible reasoning; at max it can generate an order of magnitude more tokens than the answer itself. The exact amount is a behaviour rather than a published budget, which is why any multiplier you see in a cost model is a planning assumption and not a vendor figure.
One more distinction matters. Thinking tokens are not the same as the model's parameters or its training data. They are generated at request time, they vary from call to call on identical input, and they are the single most volatile line in your bill.
Why they bill at the output rate
Billing classifies tokens by direction. Anything the provider feeds into the model is input; anything the model produces is output. Thinking is produced, so it is output.
cost = input_tokens x input_rate
+ (answer_tokens + tool_tokens + thinking_tokens) x output_rate
On Claude Sonnet 5.5 the two rates are $2 and $10 per million tokens. That is a 5x gap, so a token of thinking is five times more expensive than a token of prompt. A model that thinks for 17,100 tokens and answers in 900 has spent almost all of its money on the part you cannot read.
The same logic applies to tool calls. When a model reasons about which tool to call and what arguments to pass, those tokens are output too. Lower effort produces fewer and terser tool calls, so the saving on an agentic workload is larger than the reasoning budget alone suggests.
Why they count against max_tokens
max_tokens is a cap on everything the model generates, not just the answer. Thinking is generated, so it draws from the same budget.
That has a practical consequence. If you run at xhigh or max with a max_tokens sized for the answer alone, the model can exhaust the budget while it is still thinking. You then receive a truncated answer, or no answer at all, and you are billed for every token it produced before it stopped.
Size max_tokens for the thinking budget plus the answer, not for the answer alone. A cap that was generous at low can be the reason a max request returns nothing at all.
Visible and hidden thinking
Whether you can read the reasoning is a product setting. Whether you pay for it is not.
Some configurations stream or summarise the thinking; others hide it entirely. The billing is identical in both cases, because billing is a property of the token and not of the interface. This is the reason invisible tokens cause the most confusion: the cost appears in the usage report while nothing appears on screen.
It also means you cannot audit your spend by reading the output. A run that looks cheap because the answer was three sentences long may have generated 20,000 tokens of reasoning behind it.
Separating thinking tokens from your usage data
The usage report gives you a single output token count, and thinking is already inside it. It is not broken out as its own field, so you have to derive it.
| what you want to know | how to get it |
|---|---|
| total output billed | read the output token count from the usage report |
| answer tokens | count the answer text you received |
| thinking tokens | total output minus answer tokens |
| cost of the thinking | thinking tokens x output rate, divided by one million |
The cleanest method is a controlled difference. Send the same prompt twice, once at the effort level you care about and once at low, and subtract the output counts. The gap is what the extra effort bought you in tokens. Do it on ten representative requests rather than one, because the spread between requests is large and a single sample will mislead you.
The arithmetic on a 4,000 in / 900 out prompt
Model thinking as a multiple of the answer length — 1, 3, 8, 20 and 45 for low through max. This is a planning assumption, not a vendor figure, but it produces the right shape. Claude Sonnet 5.5 at $2 in and $10 out:
| effort | modelled thinking tokens | total output tokens | cost / call | thinking share of bill |
|---|---|---|---|---|
low | 0 | 900 | $0.0170 | 0% |
medium | 1,800 | 2,700 | $0.0350 | 51% |
high | 6,300 | 7,200 | $0.0800 | 79% |
xhigh | 17,100 | 18,000 | $0.1880 | 91% |
max | 39,600 | 40,500 | $0.4130 | 96% |
The last column is the one to remember. From high upward, thinking is not a line item, it is the bill. Trimming the prompt at max changes almost nothing, because only 1.9% of the cost is input. Trimming thinking changes everything.
Why effort is a behaviour, not a budget
Anthropic describes effort as a behavioural signal rather than a fixed token budget: at lower levels the model still thinks on hard problems, just less than it would at a higher level. That shapes how you should model cost in two ways.
First, any multiplier is an average over your traffic, not a guarantee per request. Second, the same effort setting produces very different token counts on a hard prompt and an easy one, so the monthly bill is the sum of a distribution rather than a fixed per-call price. Measure the distribution on your own traffic before you commit to a number.
Where the bill actually goes wrong
Nobody sets max by accident. The expensive surprises come from two quieter places.
The first is the default. Most Claude models default to high, and Opus 5.5 defaults to medium. Setting effort to the model default is identical to omitting the parameter, so a service that never touched effort is already paying for high.
The second is the background job: the classifier, the router, the tagger, the summariser nobody cost-reviews. They usually run at the default, they usually do not need it, and at volume they add up faster than the user-facing chat that everyone is watching.
Compare the two bills directly in the calculator: same prompt, same model, only the effort level changes.
The short version
- Thinking tokens are output tokens. They bill at the output rate, with no discount for being internal.
- They count against
max_tokens, so an answer-sized cap can truncate a high-effort request. - Visibility is a product feature; billing is not. Hidden thinking still costs money.
- The usage report folds thinking into a single output count. Derive it by differencing two effort levels.
- From
highupward, thinking is most of the bill. Attack it before you touch the prompt. - Effort is a behaviour, not a budget, so model it as a distribution rather than a fixed multiple.
Frequently asked
Are thinking tokens billed at a different rate from answer tokens?
No. They are output tokens and they bill at the output rate. There is no cheaper internal-work bucket, so a token the model spends reasoning costs exactly what a token of the answer costs.
Do thinking tokens count against max_tokens?
Yes. max_tokens caps everything the model generates, and thinking is generated. If the budget is too small at a high effort level, the model can spend it all thinking and return a short or empty answer.
Can I see the thinking tokens I am paying for?
Not always. Some settings surface the reasoning and others summarise or hide it, but billing does not change either way. Treat visibility as a product feature and billing as a separate fact.
How do I find out how many thinking tokens a request used?
Your usage report gives a single output token count that already includes them. Run the same prompt at two effort levels and subtract, or subtract the length of the answer you received from the total.