Implementation

How to run an effort sweep on your own traffic (and what to do with the results)

Sample 50 to 200 real requests, hold the model and prompt fixed, vary only the effort level, and record accuracy, latency, output tokens and cost at each setting.

Run the sweep yourself. Effort is a behavioural signal rather than a fixed token budget, and vendors recalibrate it between model generations, so the level that is right for your task cannot be read off a price list or carried over from someone else's workload. Sample 50 to 200 real requests, hold everything except effort constant, and record four numbers per level: accuracy, latency, output tokens and cost.

The output of the sweep is not a number, it is a routing table. Each traffic class gets a model and a level, and the table is what you tune the next time the bill or the quality moves.

Why sweep your own traffic

The multipliers used in most cost models, including the ones on this site, are planning assumptions. They describe the shape of the curve and nothing more. Your traffic has its own difficulty distribution, and that distribution decides how much each level actually helps.

Two workloads with identical prompts can need different levels. A summariser working on homogeneous documents may saturate at medium. A summariser working on documents of wildly varying complexity may need high on the hard ones and waste money on the easy ones at any single fixed level.

A sweep also tells you something no price table can: how often the higher level changes the answer at all. That frequency is the input to every break-even calculation, and it is the number most teams guess.

Sample 50 to 200 real requests

Pull the sample from production logs, not from a hand-written test set. The point is to capture the real distribution of difficulty, and synthetic prompts are almost always easier and more uniform than real traffic.

Fifty requests is enough to see the shape of the curve. Two hundred is enough to trust the ordering of adjacent levels, which is the decision you are actually making: is medium better than low on this task, and is high better than medium.

Stratify if you know your traffic has segments. A sample that is 90% password resets will not tell you anything about the 10% that are hard, and the hard segment is usually where the level matters.

Change one thing: effort

Hold the model, the prompt, the max_tokens and every sampling setting constant. Vary only the effort level.

If two things change at once, the result cannot tell you which one moved the metric. This is the most common way a sweep produces an unusable answer: someone upgrades the model and raises the effort in the same run, sees the accuracy improve, and cannot say which change bought it.

Run every sampled request at every level you are considering. Pairing the requests matters, because comparing low on easy requests with high on hard ones will make effort look far more powerful than it is.

One practical note: a level the model does not support will error rather than degrade, so validate the combination before the run rather than discovering it halfway through. A failed request should be recorded as a failure, not dropped, because dropping it hides the cost of the level.

Record four numbers

For each request and each level, log four things.

columnwhat it iswhy it matters
accuracyyour own pass or fail judgementthe only measure of whether the level helped
latencyp50 and p95 wall-clock timethe product cost that does not appear on the invoice
output tokensthe total generated countthe driver of both latency and cost
costoutput tokens times the output rate, plus inputthe thing you are trying to reduce

The accuracy column needs a definition before you start. A grader, a rubric, or a human review all work, but they have to be fixed in advance and applied identically at every level. If the bar moves between runs, the sweep measures the bar.

The cost column should use the same rates throughout, so the levels are comparable. Pull them from one place rather than recomputing them per run.

Plot the quality-cost frontier

Put cost on one axis and accuracy on the other, and you have a frontier. The shape is what matters: accuracy rises quickly at first, then flattens while cost keeps climbing.

Here is a result table template. The numbers below are illustrative, chosen to show the shape of a typical curve, and every one of them should be replaced with your own measurements.

effortaccuracy (yours)p95 latencyoutput tokenscost / callmonthly cost
low71%1.1 s900$0.017$1,020
medium84%2.9 s2,700$0.035$2,100
high88%7.6 s7,200$0.080$4,800
xhigh89%19 s18,000$0.188$11,280
max89%43 s40,500$0.413$24,780

The shape to look for is the flat section. In this template, high to max buys one point of accuracy for five times the money, and medium to high buys four points for 2.3 times the money. Those two steps are not comparable, and the frontier is what makes that visible.

Pick the knee, not the maximum

The knee is the point where the next level up stops buying accuracy worth paying for.

A workable rule: pick the cheapest level whose accuracy is within two points of the level above it. If two adjacent levels are within two points of each other, take the cheaper one. Then check the p95 latency against the product's budget, and if it fails, drop one level and check again.

Two points is a starting threshold, not a law. Tighten it for tasks where accuracy is the product and loosen it where errors are cheap and recoverable.

Do not pick the maximum. The maximum is the level where the budget is largest, not the level where the value is highest, and the flat section of the frontier is where the money goes to die.

Turn the result into a routing table

The sweep is only useful once it becomes configuration. Turn the knee for each traffic class into a model-and-level pair, and keep it in config rather than in code.

traffic classmodeleffortreason
classify ticketClaude Sonnet 5.5lowmechanical task, one right answer
draft replyClaude Sonnet 5.5mediumlatency matters and accuracy holds
debug a hard caseClaude Opus 5.5mediumstronger model at its own default
migration planClaude Opus 5.5xhighlong-horizon reasoning
eval harnessClaude Opus 5.5maxoffline, accuracy is the product

Note the third row. Opus 5.5 defaults to medium, so a stronger model at its default level is often the better move than a weaker model at xhigh, and it is usually cheaper as well.

Keep the routing table next to the sweep results, with the date of the run. Six months later, the date is the only way to know whether the numbers are still worth trusting.

Re-run the sweep after any model change. A new generation can shift how much a level thinks, so the knee you found on the old model may not be the cheapest acceptable setting on the new one.

Re-sweep when the model changes

Effort is calibrated per model generation. A level that produced one behaviour on Sonnet 5 can produce a different amount of thinking on Sonnet 5.5, and the same is true across an Opus version boundary.

That means the routing table has a shelf life tied to your model versions, not to a calendar. Trigger a re-sweep when you upgrade a model, when a vendor changes default behaviour, and when your own traffic mix shifts enough that the old sample no longer represents it.

The good news is that a sweep is cheap. A few hundred requests at five levels is a rounding error next to the bill it is designed to control, which makes this the highest-return measurement most teams are not running.

Run your sampled requests through the calculator to convert the token counts into a monthly figure for each level.

The short version

Frequently asked

Why not just use the recommended effort level for a task?

Because effort is a behavioural signal rather than a fixed budget, and vendors recalibrate it between model generations. A level that was right for one model can be wrong for the next, so the only reliable answer comes from your own traffic.

How many requests do I need to sample?

Fifty is enough to see the shape of the curve and two hundred is enough to trust the ordering of adjacent levels. Sample real requests rather than synthetic ones, because the distribution of difficulty is the thing you are measuring.

What should I hold constant during a sweep?

Everything except effort. Fix the model, the prompt, the max_tokens and the sampling settings, then vary only the effort level. If two things change at once, the result cannot tell you which one moved the metric.

When do I need to run the sweep again?

Whenever the model changes. A new generation can shift how much a level thinks, so the knee you found on the old model may no longer be the cheapest acceptable setting on the new one.