What a higher effort level does to latency, and how to measure it
Latency tracks output tokens rather than effort itself, so it grows with the modelled thinking budget. Measure p50 and p95, and never judge a level on time to first token alone.
Latency tracks output tokens, not effort itself. Higher effort produces more thinking tokens, output is generated one token at a time, so wall-clock time grows roughly in step with the modelled output count. On the reference shape, going from low to max multiplies the generated tokens by 45, and that is the multiplier that lands on your users' wait.
The practical consequence is that effort is not a free quality dial. Every step up buys accuracy you may not need and latency you certainly feel. If the product is interactive, latency is part of the product.
Why effort moves latency at all
A model generates output tokens sequentially. Each token is produced by a forward pass, and the next token depends on the one before it, so the work cannot be parallelised across the answer. Doubling the number of tokens roughly doubles the generation time.
Thinking tokens are output tokens, so they sit in exactly the same queue as the answer. They are generated before the answer, which is why a long reasoning pass shows up first as a long delay and then as a normal-looking answer.
There is one part of the request that does not scale with effort: the input. Prompt processing happens once, up front, and its cost depends on the prompt length rather than on the effort level. That fixed portion is why latency grows less than the output multiplier suggests.
Latency tracks output tokens, not effort
The level name is a label, not a unit. Two calls at high on the same prompt can generate very different numbers of thinking tokens, and therefore take very different amounts of time.
That is why the right input to a latency estimate is the measured output token count for the task, not the effort label. If you know that a task generates about 7,200 output tokens at high, you can estimate the time from your generation rate. If you only know the level, you cannot.
This also explains why the same effort level feels fast on one endpoint and slow on another. A classifier returning twenty tokens at high is quick. A long-form writer at high is not.
A relative wall-clock framework
Wall-clock time is roughly fixed overhead plus generated tokens divided by the generation rate. The fixed part covers network, queueing and prompt processing, and it does not move with effort.
Assume, for illustration, a fixed overhead of 0.4 seconds and a generation rate that produces the 900-token low answer in 1.0 second. The multipliers below are the same planning assumptions used elsewhere — 1, 3, 8, 20 and 45, which are not vendor figures.
| effort | relative generation time | illustrative total | vs. low |
|---|---|---|---|
low | 1.0x | 1.4 s | 1.0x |
medium | 3.0x | 3.4 s | 2.4x |
high | 8.0x | 8.4 s | 6.0x |
xhigh | 20.0x | 20.4 s | 14.6x |
max | 45.0x | 45.4 s | 32.4x |
Read the last two columns together. Generation time grows 45x, but total latency grows 32x, because the 0.4 seconds of overhead is paid once at every level. Substitute your own overhead and generation rate and the shape holds even when the absolute numbers do not.
At 2,000 calls a day
The per-call number is what a user feels. The aggregate is what the product experiences across a day.
| effort | extra wait per call vs low | cumulative extra wait, 2,000 calls / day |
|---|---|---|
low | 0 s | 0 h |
medium | 2.0 s | 1.1 h |
high | 7.0 s | 3.9 h |
xhigh | 19.0 s | 10.6 h |
max | 44.0 s | 24.4 h |
At max, the workload adds roughly a full day of cumulative waiting every day. That is the same workload that costs about $24,780 a month, so the two dials move together: the level that costs the most also makes users wait the longest.
If you run a queue rather than a synchronous call, the same numbers become throughput. A worker that takes 45 seconds per task serves far fewer tasks per hour than one that takes 1.4, so effort choices set your concurrency requirement as well as your bill.
Latency is part of the product
For a batch job, latency is an engineering detail. For anything a person is waiting on, it is the product.
A support assistant that answers in 1.4 seconds feels instant. The same assistant at 8.4 seconds feels broken, even if the answer is better. Users do not compare the answer to the one they would have got at low; they compare the wait to the wait they had last week.
This is the argument for medium as the default on customer-facing chat. It is fast enough to feel responsive and cheap enough to run at volume, and the accuracy step from medium to high rarely shows up in a conversation the user is reading in real time.
State the latency budget before you tune anything. If the product promises an answer in three seconds, then high is out of budget on this shape regardless of how good it is, and the decision collapses to the fastest level that clears the accuracy bar. A budget turns a subjective argument about quality into a constraint you can test against.
Do not tune effort on cost alone for interactive traffic. The level that saves the most money is often the level that makes the product feel slow, and that cost does not appear on the invoice.
Measure p50 and p95, not the mean
The mean hides the tail, and the tail is what generates complaints. A service with a 2-second average and a 30-second p95 will produce support tickets about the 30 seconds, not praise for the average.
| metric | what it tells you |
|---|---|
| p50 total latency | the typical experience |
| p95 total latency | the experience that generates complaints |
| time to first token | perceived responsiveness in a streaming UI |
| output tokens per call | the driver of both latency and cost |
| retries and errors | hidden latency, since a retry doubles the wait |
Record the effort level alongside each measurement. A latency change with no effort label is not actionable, because you cannot tell whether the level moved or the traffic did. Split the percentiles by task as well: a single global p95 mixes a fast classifier with a slow writer and describes neither.
Why time to first token is not enough
Time to first token measures how long before anything appears. It is the right metric for perceived responsiveness in a streaming interface, and the wrong one for total cost of waiting.
A reasoning model can spend a long time thinking before emitting anything, then stream the answer quickly. First-token time alone would call that slow, and total time would call it fast. The reverse also happens: a fast first token followed by a slow stream.
Measure both, and report them separately. If you only have one number, use total p95, because that is what the user actually waits for.
Compare the per-call cost of each level in the calculator, then multiply the latency difference by your daily volume to see what the level costs in waiting.
The short version
- Latency tracks output tokens, and output tokens scale with effort, so latency grows roughly linearly with the modelled budget.
- Total latency grows less than the output multiplier, because fixed overhead is paid once at every level.
- At 2,000 calls a day, moving from
lowtomaxadds about 24 hours of cumulative user wait per day. - For interactive traffic, latency is part of the product and belongs in the effort decision.
- Track p50 and p95, split by task, and always record the effort level next to the measurement.
- Time to first token and total time answer different questions. Measure both.
Frequently asked
Why does a higher effort level make the response slower?
Higher effort produces more thinking tokens, and output tokens are generated one at a time. More tokens to generate means more time generating them, so wall-clock latency grows roughly in step with the output count.
Does latency double when cost doubles?
Not exactly. Latency scales with output tokens, while cost includes the input floor as well. On the reference shape, output tokens grow about 24x from low to max but total latency grows less than that because fixed overhead and input handling do not scale.
Should I measure average latency or percentiles?
Percentiles. The mean hides the tail, and the tail is what users complain about. Track p50 for the typical experience and p95 for the experience that generates support tickets.
Is time to first token enough to judge responsiveness?
No. A streaming interface can show a fast first token and then stall, and a reasoning model can have a long first-token delay followed by fast output. Measure first-token time and total time together.