Guides
Long-form notes on Claude effort levels, thinking-token billing and what an API bill actually consists of. Every figure is worked through from the same price table the calculator uses.
Implementation
- Setting the effort level in the Claude API: defaults, errors and worked examples The effort level lives in output_config.effort, defaults to high on most models and medium on Opus 5.5, and errors instead of silently downgrading when a combination is unsupported.
- How to run an effort sweep on your own traffic (and what to do with the results) Sample 50 to 200 real requests, hold the model and prompt fixed, vary only the effort level, and record accuracy, latency, output tokens and cost at each setting.
Choosing a level
- What a higher effort level does to latency, and how to measure it Latency tracks output tokens rather than effort itself, so it grows with the modelled thinking budget. Measure p50 and p95, and never judge a level on time to first token alone.
- When Claude's max effort level is actually worth the money Max costs about $0.333 more per call than high on the reference shape. It pays only when a single wrong answer is worth more than that, multiplied by how often it prevents one.
Cost
- Prompt caching vs raising effort: which lever actually saves more money Caching is free quality-wise but small on output-heavy traffic, while effort is the larger lever. To save $1,000 a month you need either about 9,300 cacheable tokens or a single effort step.
- What the Claude effort parameter actually charges you for One parameter, an 11x swing on the same prompt. Where the money actually goes, and the counterintuitive result that a bigger model at moderate effort can beat a smaller model at extreme effort.
- Seven levers that cut a Claude API bill, ranked by how much they save Matching effort to traffic class is the largest lever on output-heavy work, caching is the safest, and the rest range from a few hundred to a few thousand dollars a month.