Implementation

Setting the effort level in the Claude API: defaults, errors and worked examples

The effort level lives in output_config.effort, defaults to high on most models and medium on Opus 5.5, and errors instead of silently downgrading when a combination is unsupported.

The effort level is set through output_config.effort and takes five values: low, medium, high, xhigh and max. Two things cause most of the damage. The default is high rather than off — and medium on Opus 5.5 — so a service that never sets it is already paying for reasoning. And an unsupported combination raises an error rather than degrading quietly, which means a batch job can fail on a request you thought was safe.

Everything below is the practical version: the values, the defaults, two ways to set it, and the error case worth planning for.

The five values

levelwhat it is fortypical traffic
lowDeterministic work with one right answerclassify, extract, tag, route
mediumFast, natural output where latency showsinteractive chat, summaries
highThe general-purpose defaultcode, analysis, most production traffic
xhighLong-horizon reasoning across many stepsmigrations, architecture, planning
maxThe largest thinking budget availablebenchmark-grade problems, rare in production

OpenAI's equivalent parameter is reasoning_effort, which accepts minimal, low, medium and high. GPT-6 Astra supports only low, medium and high, so a port between providers is not a one-to-one mapping of level names.

The names are labels rather than a numeric scale. There is no published token budget behind each one, and the same label can produce different amounts of thinking on different models. That is why a level cannot be copied between providers or between model generations without re-measuring, and why the cost of a level is best read from your own token counts rather than from the name.

The default trap

The default is not "off". It is a level, and a fairly expensive one.

modeldefault effort
Claude Sonnet 5.5high
Claude Sonnet 5high
Claude Opus 5.5medium
GPT-6 Astraset through reasoning_effort

Setting effort to the model default is identical to omitting the parameter, so there is no behavioural difference between the two. What that means in practice is that a classifier written before anyone thought about cost is running at high, and nobody notices until the invoice arrives.

The Opus 5.5 case is worth internalising for a different reason. If you move a task from Sonnet 5.5 to Opus 5.5 and set effort to high to be consistent, you have changed two things at once. Opus 5.5's own default is medium, so high on Opus is a deliberate step up rather than a like-for-like move.

Top-level effort

The top-level form sets one effort level for the whole request.

{
  "model": "claude-sonnet-5-5",
  "max_tokens": 4096,
  "output_config": { "effort": "high" },
  "messages": [
    { "role": "user", "content": "Summarise this incident report." }
  ]
}

This is the form to use when the whole request is one kind of work. It is simple, it is easy to log, and it is easy to put behind a config key.

It is also the form that breaks prompt caching when you vary it. Effort shapes the rendered prompt, so changing it between requests drops the cached prefix and you pay a cache write instead of a read.

Per-message effort

The per-message form sets effort on a single message, leaving everything before it untouched.

{
  "model": "claude-sonnet-5-5",
  "max_tokens": 4096,
  "messages": [
    { "role": "user", "content": "<long stable system context>" },
    { "role": "user", "content": "Extract the dates.",
      "output_config": { "effort": "low" } },
    { "role": "user", "content": "Now explain the timeline.",
      "output_config": { "effort": "high" } }
  ]
}

The nesting mirrors the top-level field, and the important property is positional. Because the earlier messages are unchanged, the cacheable prefix stays valid and only the new work is affected by the effort setting.

Use this form when one conversation contains both cheap and expensive steps: an extraction at low followed by a reasoning step at high, inside a single cached context.

Where the setting lives in the call path

There are three places the value can come from, and mixing them without a rule is how effort ends up inconsistent across a codebase.

Put the default in configuration, keyed by task name. Put a per-request override on the call for the handful of paths that genuinely differ. Put the per-message override inside the one conversation that needs mixed levels, and nowhere else.

The rule to hold to is that exactly one layer decides. If the config says medium and a client hardcodes high, the hardcode wins silently and the config becomes a lie. That is how a tuning change gets rolled out to production and appears to do nothing.

Log the resolved level on every request, next to the token counts. A log line with the effort level, the input tokens and the output tokens is enough to explain almost any bill movement after the fact, and it costs nothing to write.

When the request errors instead of degrading

The API does not quietly fall back to a supported level when you ask for something it cannot do. It returns an error, and the request fails.

Sonnet 5.5 has one case that catches people out: at xhigh and max it refuses the minimal setting that would disable thinking entirely, and returns an error rather than silently turning thinking back on. The message is:

This model does not support enabling thinking with effort=xhigh

The design choice behind that is worth understanding. A silent downgrade would make your cost model wrong without telling you. An error makes it your problem, which is annoying once and correct forever after.

Treat any effort change as a request-shape change. Test it against the real model before you roll it out, and log the error body rather than the status code alone.

Configure it, do not hardcode it

Effort is the highest-leverage cost setting in the request, and you will change it more than once. Hardcoding it in the client means a deploy for every tuning pass.

Keep a mapping from task to effort in configuration:

Then log the resolved level alongside the token counts on every request. Without that log you cannot tell whether a bill moved because of effort or because of traffic, and you will be guessing on the next tuning pass.

Batch jobs need a fallback path

Batch work is where the error case becomes expensive. If you submit ten thousand requests at a level the model does not support, you do not lose one request, you lose the batch.

Two things make that survivable. Validate the level against the model before you submit, and handle the error per request rather than per batch, so a single bad request does not take the rest down with it. A fallback to the nearest supported level is better than a failed job, as long as you record that the fallback happened.

Work out which level each task should use with the calculator, then read the per-call cost back into the config map above.

The short version

Frequently asked

What is the default effort level if I never set it?

Most Claude models default to high, and Opus 5.5 defaults to medium. Setting effort to the model default is identical to omitting the parameter, so leaving it out does not mean it is off.

How do I set effort per message instead of per request?

Put the effort setting on the individual message rather than on the top-level request. This keeps everything before that message identical, so a cached prefix stays valid.

What happens if I request an unsupported effort combination?

The request errors rather than quietly downgrading. Sonnet 5.5, for example, rejects the minimal thinking-off setting at xhigh and max with the message This model does not support enabling thinking with effort=xhigh.

Should I hardcode the effort level in the request?

No. Put it in configuration keyed by task so you can change it without a deploy. Effort is the single highest-leverage cost setting you have, and you will want to tune it more than once.