What the Claude effort parameter actually charges you for
One parameter, an 11x swing on the same prompt. Where the money actually goes, and the counterintuitive result that a bigger model at moderate effort can beat a smaller model at extreme effort.
Take a support bot on Claude Sonnet 5.5. About 4,000 input tokens of context, about 900 output tokens per reply, 2,000 calls a day. Now change one string in the request body:
output_config.effort: "low" -> "xhigh"
Same model. Same prompt. Same 900-token answer.
The bill goes from roughly $1,020 a month to roughly $11,280 a month.
That is an 11x swing from a parameter most teams never touch, and the mechanism behind it is not the one most people assume. Once you see where the money goes, you stop reaching for max out of reflex.
What the parameter actually controls
Claude models accept an effort setting through output_config.effort, with five levels: low, medium, high, xhigh and max. OpenAI's equivalent is reasoning_effort, which takes minimal, low, medium or high.
The one-line description — "how much it thinks" — is too narrow. Effort affects every token in the response: the answer text, tool calls and their arguments, and thinking when it is active. It is a how-much-work dial, not a thinking-token dial. Lower effort also produces fewer and terser tool calls, which is easy to miss if you are only watching the reasoning budget.
It is also a behavioural signal rather than a strict budget. Anthropic's documentation is explicit that at lower effort levels the model still thinks on sufficiently difficult problems, just less than it would at a higher level for the same problem. Hold on to that, because it is the reason the multipliers further down are labelled assumptions.
Why any of this reaches your bill: thinking tokens are output tokens. They are generated by the model and they count against max_tokens alongside the answer text, so they bill at the output rate — the expensive column. Cost per call is roughly:
cost = input_tokens x input_rate
+ (answer_tokens + tool_tokens + thinking_tokens) x output_rate
Effort moves everything inside the second bracket. That is the whole mechanism.
The arithmetic, worked through
To make this concrete, model effort as scaling the output side by a fixed factor — 1, 3, 8, 20 and 45 for low through max. That is a planning simplification, not a published figure, but it is enough to show the shape of the problem.
Using Claude Sonnet 5.5 at $2 per million input tokens and $10 per million output tokens, with the same 4,000 in / 900 out shape:
| effort | modelled output tokens | cost / call | vs. low | 2,000 calls/day |
|---|---|---|---|---|
low | 900 | $0.0170 | 1.0x | $1,020 / mo |
medium | 2,700 | $0.0350 | 2.1x | $2,100 / mo |
high | 7,200 | $0.0800 | 4.7x | $4,800 / mo |
xhigh | 18,000 | $0.1880 | 11.1x | $11,280 / mo |
max | 40,500 | $0.4130 | 24.3x | $24,780 / mo |
Read the second column against the third. You asked for a 900-token answer, and at max the model is billing you for roughly 40,500 tokens of work to produce it.
Why 45x becomes 24x
The modelled output multiplier from low to max is 45x. The multiplier on your actual bill is 24.3x. The gap is the input side: those 4,000 input tokens cost the same $0.008 whether you are at low or max. Input is a floor under your bill, and it does not move.
That floor matters more than most people expect:
| effort | input share of total cost |
|---|---|
low | 47.1% |
medium | 22.9% |
high | 10.0% |
xhigh | 4.3% |
max | 1.9% |
If you are running at low or medium, almost half your bill is input tokens. Optimising the output side there is nearly pointless — trim the prompt, drop context you do not need, or cache the static part. At xhigh the same prompt work buys you almost nothing, and the thinking budget is the thing to attack instead.
Most cost-optimisation advice is written by people running at high or above. It is simply wrong for the low and medium case, which is where a large share of production traffic actually lives.
Switch model before raising effort
Here is the comparison that changes the decision. Same 4,000 in / 900 out shape, everything at high:
| model | input / output rate | cost / call | 2,000 calls/day |
|---|---|---|---|
| Claude Sonnet 5.5 | $2 / $10 | $0.0800 | $4,800 / mo |
| Claude Opus 5.5 | $4 / $20 | $0.1600 | $9,600 / mo |
| GPT-6 Astra | $10 / $50 | $0.4000 | $24,000 / mo |
Now the two lines worth staring at:
Claude Opus 5.5 at high costs about $0.16 per call. Claude Sonnet 5.5 at xhigh costs about $0.188 per call.
The stronger model at moderate effort is cheaper than the weaker model at extreme effort. That is not a rounding artefact — it falls straight out of the two rate pairs, and it inverts the instinct most of us have. When an answer is bad, we crank the effort dial. Often the better move is to change which model is thinking, not how hard it is thinking.
The two settings are not interchangeable, so this is a hypothesis to test against your own evals rather than a rule. But it is a cheap hypothesis, and it is the first thing worth trying.
Two footnotes worth knowing before you ship
1. Changing effort at the top level invalidates your prompt cache. Effort shapes the rendered prompt, so flipping it between requests drops the cached prefix. If you use prompt caching and want to vary effort per turn, use per-message effort instead — everything before that message stays cached.
2. The default is not what you would guess. high is the default on most models, but Opus 5.5 defaults to medium. Setting effort to the model default is identical to omitting the parameter entirely, so "I never set it" does not mean "it is off" — it means you are already paying for high, or for medium on Opus 5.5.
That second point is the one that costs the most money in practice. Nobody sets max by accident. Plenty of teams never realise their classifier is running at high.
The order to try things in
The tables above suggest a sequence, and it is not the sequence most teams follow.
First, ask whether the task needs reasoning at all. Classification, extraction, formatting and translation have one correct shape of answer. If your traffic is mostly that, the answer is low on the cheapest model that passes your evals, and there is nothing else to optimise.
Second, if you are already at low or medium, work on the input side. At those levels roughly a third to a half of your bill is input tokens. Trimming retrieved context, dropping a few-shot example nobody uses, or caching a stable system prompt moves the number. Raising or lowering effort does not, because the output side is already the smaller half.
Third, if you are at high or above and quality is the problem, compare a model swap against an effort bump. A stronger model at high can be cheaper than a weaker model at xhigh, and it is usually better at the thing you were trying to fix.
Fourth, only then spend on max. It is a tool for a handful of requests where a wrong answer is genuinely expensive, not a default for a queue.
What to default to, by task
| task | setting | why |
|---|---|---|
| Classify, extract, tag | low, cheapest model | Deterministic output. Paying for reasoning here is pure waste. |
| Translate or rewrite | low | Modern models are strong enough that effort buys almost nothing. |
| Customer-facing chat | medium | Latency is part of the product. Medium is the sweet spot. |
| Write or refactor code | high | The default workhorse. Most production traffic belongs here. |
| Debug a hard production issue | high on the flagship | Once a bug is expensive, model quality beats token price. |
| Architecture or migration | xhigh | Long-horizon planning is what the level exists for. |
| Benchmark-grade problems | max | Rarely worth it in production. |
Two rules fall out of the tables above. Below high, attack the input side — the thinking budget is not your problem yet. Above high, attack the thinking budget, and first ask whether a better model at lower effort is cheaper.
The short version
- Effort scales every output token — answer text, tool calls and thinking — and they all bill at the output rate.
- The multiplier on output tokens (45x) is not the multiplier on your bill (24x), because input cost is a fixed floor.
- At
lowandmedium, roughly half your bill is input. Optimise the prompt, not the output. - A stronger model at moderate effort can be cheaper than a weaker model at extreme effort. Test that before you turn the dial.
- The default is not "off". It is
highon most models.
Check your own numbers. The calculator takes your input tokens, output tokens and call volume, and prints the per-call and monthly cost at every level and every model — with the multipliers shown underneath so you can see the assumptions rather than trust them. It runs entirely in your browser.
Frequently asked
Does the effort level change my input token cost?
No. Input tokens are billed at the same rate regardless of effort. Effort scales the output side only, which is why the multiple on your total bill is always smaller than the multiple on your output tokens.
Is a stronger model at lower effort really cheaper than a weaker model at higher effort?
It can be, and the arithmetic is easy to check. At a 4,000 in / 900 out shape, Claude Opus 5.5 at high costs about $0.16 per call, while Claude Sonnet 5.5 at xhigh costs about $0.188. The two settings are not equivalent, so treat it as a hypothesis to test against your own evals rather than a rule.
What is the default effort level if I never set it?
Most Claude models default to high; Opus 5.5 defaults to medium. Setting effort to the model default is identical to omitting the parameter, so not setting it does not mean it is off.
Do the thinking-token multipliers come from Anthropic?
No. The 1 / 3 / 8 / 20 / 45 multipliers used here are planning assumptions, not published vendor figures. Anthropic describes effort as a behavioural signal rather than a strict token budget. Measure your own traffic and replace them.