Prompt caching vs raising effort: which lever actually saves more money
Caching is free quality-wise but small on output-heavy traffic, while effort is the larger lever. To save $1,000 a month you need either about 9,300 cacheable tokens or a single effort step.
If your bill is too high, the larger lever is almost always effort, not caching. On the reference shape — 4,000 in, 900 out, Claude Sonnet 5.5 — caching the stable prefix saves roughly $324 a month at high, while dropping from high to low saves roughly $3,780. But caching costs nothing in quality, so the order is not a choice: cache first, then move effort.
The two levers move different sides of the bill, which is why the comparison is shape-dependent rather than absolute.
What each lever actually moves
Caching moves the input side. Effort moves the output side. On the reference shape the input side is only 10% of the bill at high, so caching has a low ceiling there. Flip the shape to a long-context retrieval task and the ranking reverses, because the input portion becomes most of the bill.
| lever | side of the bill it moves | what changes | ceiling |
|---|---|---|---|
| prompt caching | input | $2.00 down to $0.20 per million on cached tokens | the whole input portion |
| lower effort | output | 1x to 45x modelled thinking tokens | most of the output portion |
The ceiling is the thing to estimate before you optimise. If input is 10% of your bill, no amount of caching saves more than 10%. If output is 90%, effort is where the money is.
What prompt caching saves
On Claude Sonnet 5.5 the three relevant rates are $2 per million for fresh input, $0.20 per million for a cache read, and $2.50 per million for a cache write. A read is 90% cheaper than fresh input. A write is 25% more expensive than fresh input.
Take the reference shape with a stable 3,000-token prefix reused on every call, and 1,000 tokens of genuinely new input:
- Without caching: 4,000 x $2 per million = $0.0080 per call.
- With a cache hit: 3,000 x $0.20 per million = $0.0006, plus 1,000 x $2 per million = $0.0020, so $0.0026 per call.
That is a saving of $0.0054 per call, or about 67% of the input cost. At 2,000 calls a day over 30 days, 60,000 calls, the monthly saving is about $324.
Notice what caching does not do: it does not change the answer. The model sees the same tokens, rendered the same way, at a lower price. That makes it the safest lever on the list, and the first one to pull.
What lowering effort saves
Now run the same arithmetic on the output side, holding the prompt fixed. From the reference shape at Claude Sonnet 5.5 rates:
| effort | cost / call | 60,000 calls / month | saving vs high |
|---|---|---|---|
low | $0.0170 | $1,020 | $3,780 |
medium | $0.0350 | $2,100 | $2,700 |
high | $0.0800 | $4,800 | baseline |
xhigh | $0.1880 | $11,280 | minus $6,480 |
max | $0.4130 | $24,780 | minus $19,980 |
Dropping from high to low saves $3,780 a month on the same 60,000 calls. That is about 11.7x the caching saving on the identical shape, and it comes from one string in the request body.
The catch is real, though. Caching is free quality-wise and effort is not. Lowering effort changes the answer, so it has to be validated against your own evals before it goes to production. That is the whole trade: a larger saving for a cost that is not measured in dollars.
Saving $1,000 a month: two routes
Suppose you need to take $1,000 a month out of this workload. Over 60,000 calls that is $0.01667 per call. Here is what each route has to do to get there.
| route | what it takes | quality cost |
|---|---|---|
| cache the stable prefix | about 9,300 tokens of reusable prefix at a 90% hit rate | none, the output is identical |
drop effort from high to medium | $0.045 saved per call, clearing the target 2.7x over | the answer changes |
drop effort from high to low | $0.063 saved per call, clearing the target 3.8x over | the answer changes more |
The cache figure comes from the read discount: each cached token saves $1.80 per million, so $1,000 over 60,000 calls needs about 9,300 cached tokens. That is a large prefix. A single effort step does the same job with no prompt engineering at all, provided the quality holds.
If you have both a large prefix and slack in the effort level, you do not have to choose. Take the caching saving, then take the effort saving on top.
Cache write is a premium, not a discount
It is easy to treat caching as uniformly cheaper. It is not. The write costs more than fresh input.
- Two uncached calls on a million tokens of prefix: 2 x $2.00 = $4.00.
- One write plus one read: $2.50 + $0.20 = $2.70.
Caching therefore pays from the second reuse onward. If the prefix changes on every request, every call is a write and you pay 25% more than you would have paid without caching at all. A prefix that is rebuilt per request is the one case where caching is a straight loss.
Where the two levers interfere
The interference is a single mechanism: effort shapes the rendered prompt. Change it at the top level and the cached prefix no longer matches, so the next request is a cache miss and you pay a write again.
If you vary effort per request at the top level, you turn a workload with a 90% hit rate into a workload with a 0% hit rate. You then pay the write premium on every call, and the caching lever does nothing.
The fix is per-message effort. Setting effort on a message rather than on the request keeps everything before that message identical, so the prefix stays cached and only the new work is affected.
Do not combine top-level effort changes with prompt caching. Either fix the effort level for the request, or move the variation down to per-message effort.
When caching is the bigger lever
The ranking reverses on input-heavy traffic, which is why the shape matters more than the lever. Take a retrieval task with 100,000 input tokens and a 500-token answer. Input at $2 per million is $0.20 per call. Output at high, with the modelled 8x multiplier, is 4,000 tokens at $10 per million, or $0.04. Input is 83% of the bill, and effort can only reach the other 17%.
| lever | saving per call | what it touches |
|---|---|---|
| cache a 90,000-token prefix | $0.162 | the input side |
drop high to low | $0.035 | the output side |
Caching the 90,000-token prefix cuts it from $0.18 to $0.018 and leaves $0.02 of fresh input, so the input side falls from $0.20 to $0.038. The effort lever, on the same call, can only touch the $0.04 output portion. Caching is roughly 4.6x the saving here, the mirror image of the reference shape.
The rule that falls out is simple: look at the input share of your bill before you decide which lever to pull. Above roughly 50%, cache. Below it, move effort.
The decision rule
- Cache any prefix that is reused two or more times. It is free quality-wise, so there is no reason not to.
- Do not cache a prefix that changes on every request. You would be paying the write premium for nothing.
- Once caching is in place, move effort. That is where the large multiples live.
- Never vary top-level effort on a cached prefix. Use per-message effort if you need to vary it at all.
- Re-check the ranking whenever the shape changes. Input-heavy traffic and output-heavy traffic have opposite answers.
Model both levers side by side in the calculator: it prints the per-call and monthly cost for every effort level, and the input share tells you which lever has room to move.
The short version
- Caching moves the input side; effort moves the output side. Effort is bigger on output-heavy traffic.
- On the reference shape, caching saves about $324 a month and dropping to
lowsaves about $3,780. - A cache read is 90% cheaper than fresh input on Sonnet 5.5; a cache write is 25% more expensive.
- Caching pays off from the second reuse onward. A prefix rebuilt every call is a loss.
- To save $1,000 a month here you need roughly 9,300 cacheable tokens, but only one effort step.
- Top-level effort changes invalidate the cache. Use per-message effort instead.
Frequently asked
Which saves more, prompt caching or lowering the effort level?
On output-heavy traffic, lowering effort by far. On the reference shape, dropping from high to low saves about $3,780 a month while caching the stable prefix saves about $324. Caching wins only when the input side dominates the bill.
Is caching worth it if the prefix is only reused twice?
Yes, and barely. Two uncached calls cost $4 per million tokens of prefix, while a cache write plus a read costs $2.70, so caching pays from the second reuse onward. Anything less than a second reuse is a loss.
Why does my cache stop working when I change the effort level?
Effort shapes the rendered prompt, so changing it at the top level invalidates the cached prefix. Use per-message effort instead, which leaves everything before that message cached.
Can I use both levers at once?
Yes, and you should. Cache the stable prefix first because it costs nothing in quality, then move effort on the traffic that does not need the higher level. Just never vary top-level effort on a cached prefix.