Choosing a level

When Claude's max effort level is actually worth the money

Max costs about $0.333 more per call than high on the reference shape. It pays only when a single wrong answer is worth more than that, multiplied by how often it prevents one.

Max is worth it in one situation: when a single wrong answer costs more than the extra tokens it takes to avoid it. On the reference shape — 4,000 in, 900 out, Claude Sonnet 5.5 — moving from high to max adds about $0.333 per call. If a wrong answer costs $50 to fix, max needs to prevent one error in every 150 calls to break even. If it costs $5,000, it needs to prevent one in 15,000.

That is the whole decision. Everything else is a matter of measuring how often max actually changes the answer, which is the number nobody has until they test it.

Max is a 45x thinking budget, not a 45x bill

The modelled output multiplier from low to max is 45x. The multiplier on the bill is about 24x. The gap is the input side, which costs the same at every level.

effortmodelled output tokenscost / callvs. low2,000 calls / day
low900$0.01701.0x$1,020 / mo
medium2,700$0.03502.1x$2,100 / mo
high7,200$0.08004.7x$4,800 / mo
xhigh18,000$0.188011.1x$11,280 / mo
max40,500$0.413024.3x$24,780 / mo

Two numbers matter for the decision. The step from high to max adds $0.333 per call, or $666 a day at 2,000 calls. The step from high to xhigh adds $0.108 per call, which is a third of that for a level that often captures most of the gain.

The 45x figure is a planning assumption, not a published vendor number. Treat it as the shape of the curve rather than a guarantee, and replace it with your own measurements as soon as you have them.

Quality saturates before the budget does

The cost curve is steep and roughly linear in the modelled thinking budget. The quality curve is not.

Effort is a behavioural signal rather than a fixed budget: at a higher level the model thinks more, but the return on that thinking falls as the problem gets easier. A hard problem may benefit from the last step of effort. A routine one will not, and most production traffic is routine.

This is why the modelled 45x budget and the real accuracy gain are not the same multiple. The budget keeps growing; the accuracy flattens. The gap between those two curves is where money is wasted, and it is widest exactly at the top.

The threshold: what is one wrong answer worth

The break-even calculation needs three numbers: the extra cost per call, the cost of one wrong answer, and how often the higher level prevents one.

cost of one wrong answerextra cost per callbreak-even rate
$5$0.333max must prevent 1 error in 15 calls
$50$0.3331 error in 150 calls
$500$0.3331 error in 1,500 calls
$5,000$0.3331 error in 15,000 calls

Read the table from the right column. If a wrong answer costs $5, you need max to fix one call in fifteen to justify itself, which is a very high bar for a model that is already good. If a wrong answer costs $5,000, one fixed call in fifteen thousand pays for the whole change.

The honest problem is that the break-even rate is the one number you cannot read off a price list. It depends on your eval, and it is the reason an effort sweep is worth running before you commit.

The middle option is usually enough

The comparison that most often kills the case for max is the step immediately below it. On the reference shape, high to xhigh adds $0.108 per call, while high to max adds $0.333. The cheaper step is about a third of the price.

stepextra cost per callextra cost per monthwhat it buys
high to xhigh$0.108$6,480most of the deep-reasoning gain
xhigh to max$0.225$13,500the smallest increment
high to max$0.333$19,980the two steps combined

Read the second and third rows together. The final step costs more than twice the first one and buys the least. That is the flat section of the quality curve, and it is where the money goes if you reach for the top level by default.

So the practical sequence is: test xhigh before max, and only pay for max if you have measured that xhigh still fails on the cases that matter. On most production traffic the answer will be that it does not.

Keep the xhigh result in the routing table either way. If xhigh and max score within a point of each other on your eval, the cheaper level is the correct answer, and the comparison is the evidence you will want the next time someone asks why the top setting is not in use.

When max is worth it

When max is not worth it

Max is not a quality setting you leave on. It is a targeted spend that needs a measured justification, and the justification is a number of dollars per prevented error.

Why benchmarks and production give opposite answers

On a benchmark, every problem is graded and a wrong answer is a total loss. The score is the only thing that matters, there is no latency budget, and there is no cheaper model to fall back to. Under those rules, max is close to free: the marginal cost of the last few points is irrelevant next to the cost of being wrong.

Production inverts every one of those assumptions. Most traffic is routine, so the marginal answer does not improve. Errors are often caught by a downstream check, a human review, or a retry, so a wrong answer is not a total loss. Latency is a real product cost, and the volume is high enough that a per-call delta becomes a line item.

So the same model, on the same task, can be correctly run at max on the eval set and correctly run at medium in production. The difference is not the model. It is what a wrong answer costs.

Put your own numbers into the calculator to see the per-call delta at your volume before you decide.

The short version

Frequently asked

Does max effort cost 45x more than low?

No. The modelled thinking budget grows 45x, but the bill grows about 24x, because input tokens cost the same at every level and act as a floor. On the reference shape, low is $0.017 per call and max is $0.413.

How do I decide whether max is worth it?

Compare the extra cost per call with the cost of one wrong answer. Moving from high to max adds about $0.333 per call, so max needs to prevent one error per 150 calls if an error costs $50, or one per 15,000 if it costs $5,000.

Why do benchmarks favour max but production does not?

A benchmark grades every problem and a wrong answer is a total loss, so the score is the only cost that matters. Production traffic is mostly routine, errors are often caught by other layers, and latency is a real cost.

Does accuracy keep improving as effort rises?

No. Quality tends to plateau well before the budget does, so the last steps of effort buy less than the first ones. Measure the plateau on your own traffic rather than assuming it continues.