How to run an effort sweep on your own traffic (and what to do with the results)
Sample 50 to 200 real requests, hold the model and prompt fixed, vary only the effort level, and record accuracy, latency, output tokens and cost at each setting.
Run the sweep yourself. Effort is a behavioural signal rather than a fixed token budget, and vendors recalibrate it between model generations, so the level that is right for your task cannot be read off a price list or carried over from someone else's workload. Sample 50 to 200 real requests, hold everything except effort constant, and record four numbers per level: accuracy, latency, output tokens and cost.
The output of the sweep is not a number, it is a routing table. Each traffic class gets a model and a level, and the table is what you tune the next time the bill or the quality moves.
Why sweep your own traffic
The multipliers used in most cost models, including the ones on this site, are planning assumptions. They describe the shape of the curve and nothing more. Your traffic has its own difficulty distribution, and that distribution decides how much each level actually helps.
Two workloads with identical prompts can need different levels. A summariser working on homogeneous documents may saturate at medium. A summariser working on documents of wildly varying complexity may need high on the hard ones and waste money on the easy ones at any single fixed level.
A sweep also tells you something no price table can: how often the higher level changes the answer at all. That frequency is the input to every break-even calculation, and it is the number most teams guess.
Sample 50 to 200 real requests
Pull the sample from production logs, not from a hand-written test set. The point is to capture the real distribution of difficulty, and synthetic prompts are almost always easier and more uniform than real traffic.
Fifty requests is enough to see the shape of the curve. Two hundred is enough to trust the ordering of adjacent levels, which is the decision you are actually making: is medium better than low on this task, and is high better than medium.
Stratify if you know your traffic has segments. A sample that is 90% password resets will not tell you anything about the 10% that are hard, and the hard segment is usually where the level matters.
Change one thing: effort
Hold the model, the prompt, the max_tokens and every sampling setting constant. Vary only the effort level.
If two things change at once, the result cannot tell you which one moved the metric. This is the most common way a sweep produces an unusable answer: someone upgrades the model and raises the effort in the same run, sees the accuracy improve, and cannot say which change bought it.
Run every sampled request at every level you are considering. Pairing the requests matters, because comparing low on easy requests with high on hard ones will make effort look far more powerful than it is.
One practical note: a level the model does not support will error rather than degrade, so validate the combination before the run rather than discovering it halfway through. A failed request should be recorded as a failure, not dropped, because dropping it hides the cost of the level.
Record four numbers
For each request and each level, log four things.
| column | what it is | why it matters |
|---|---|---|
| accuracy | your own pass or fail judgement | the only measure of whether the level helped |
| latency | p50 and p95 wall-clock time | the product cost that does not appear on the invoice |
| output tokens | the total generated count | the driver of both latency and cost |
| cost | output tokens times the output rate, plus input | the thing you are trying to reduce |
The accuracy column needs a definition before you start. A grader, a rubric, or a human review all work, but they have to be fixed in advance and applied identically at every level. If the bar moves between runs, the sweep measures the bar.
The cost column should use the same rates throughout, so the levels are comparable. Pull them from one place rather than recomputing them per run.
Plot the quality-cost frontier
Put cost on one axis and accuracy on the other, and you have a frontier. The shape is what matters: accuracy rises quickly at first, then flattens while cost keeps climbing.
Here is a result table template. The numbers below are illustrative, chosen to show the shape of a typical curve, and every one of them should be replaced with your own measurements.
| effort | accuracy (yours) | p95 latency | output tokens | cost / call | monthly cost |
|---|---|---|---|---|---|
low | 71% | 1.1 s | 900 | $0.017 | $1,020 |
medium | 84% | 2.9 s | 2,700 | $0.035 | $2,100 |
high | 88% | 7.6 s | 7,200 | $0.080 | $4,800 |
xhigh | 89% | 19 s | 18,000 | $0.188 | $11,280 |
max | 89% | 43 s | 40,500 | $0.413 | $24,780 |
The shape to look for is the flat section. In this template, high to max buys one point of accuracy for five times the money, and medium to high buys four points for 2.3 times the money. Those two steps are not comparable, and the frontier is what makes that visible.
Pick the knee, not the maximum
The knee is the point where the next level up stops buying accuracy worth paying for.
A workable rule: pick the cheapest level whose accuracy is within two points of the level above it. If two adjacent levels are within two points of each other, take the cheaper one. Then check the p95 latency against the product's budget, and if it fails, drop one level and check again.
Two points is a starting threshold, not a law. Tighten it for tasks where accuracy is the product and loosen it where errors are cheap and recoverable.
Do not pick the maximum. The maximum is the level where the budget is largest, not the level where the value is highest, and the flat section of the frontier is where the money goes to die.
Turn the result into a routing table
The sweep is only useful once it becomes configuration. Turn the knee for each traffic class into a model-and-level pair, and keep it in config rather than in code.
| traffic class | model | effort | reason |
|---|---|---|---|
| classify ticket | Claude Sonnet 5.5 | low | mechanical task, one right answer |
| draft reply | Claude Sonnet 5.5 | medium | latency matters and accuracy holds |
| debug a hard case | Claude Opus 5.5 | medium | stronger model at its own default |
| migration plan | Claude Opus 5.5 | xhigh | long-horizon reasoning |
| eval harness | Claude Opus 5.5 | max | offline, accuracy is the product |
Note the third row. Opus 5.5 defaults to medium, so a stronger model at its default level is often the better move than a weaker model at xhigh, and it is usually cheaper as well.
Keep the routing table next to the sweep results, with the date of the run. Six months later, the date is the only way to know whether the numbers are still worth trusting.
Re-run the sweep after any model change. A new generation can shift how much a level thinks, so the knee you found on the old model may not be the cheapest acceptable setting on the new one.
Re-sweep when the model changes
Effort is calibrated per model generation. A level that produced one behaviour on Sonnet 5 can produce a different amount of thinking on Sonnet 5.5, and the same is true across an Opus version boundary.
That means the routing table has a shelf life tied to your model versions, not to a calendar. Trigger a re-sweep when you upgrade a model, when a vendor changes default behaviour, and when your own traffic mix shifts enough that the old sample no longer represents it.
The good news is that a sweep is cheap. A few hundred requests at five levels is a rounding error next to the bill it is designed to control, which makes this the highest-return measurement most teams are not running.
Run your sampled requests through the calculator to convert the token counts into a monthly figure for each level.
The short version
- Effort is a behaviour, not a fixed budget, so the right level for your task has to be measured on your own traffic.
- Sample 50 to 200 real requests, and stratify if your traffic has a hard segment.
- Hold the model, prompt and settings constant and vary only effort, or the result is uninterpretable.
- Record accuracy, latency, output tokens and cost for every request at every level.
- Pick the knee: the cheapest level within about two points of the level above it, subject to the latency budget.
- Turn the knee into a routing table in configuration, and re-sweep whenever the model changes.
Frequently asked
Why not just use the recommended effort level for a task?
Because effort is a behavioural signal rather than a fixed budget, and vendors recalibrate it between model generations. A level that was right for one model can be wrong for the next, so the only reliable answer comes from your own traffic.
How many requests do I need to sample?
Fifty is enough to see the shape of the curve and two hundred is enough to trust the ordering of adjacent levels. Sample real requests rather than synthetic ones, because the distribution of difficulty is the thing you are measuring.
What should I hold constant during a sweep?
Everything except effort. Fix the model, the prompt, the max_tokens and the sampling settings, then vary only the effort level. If two things change at once, the result cannot tell you which one moved the metric.
When do I need to run the sweep again?
Whenever the model changes. A new generation can shift how much a level thinks, so the knee you found on the old model may no longer be the cheapest acceptable setting on the new one.