Cost

Prompt caching vs raising effort: which lever actually saves more money

Caching is free quality-wise but small on output-heavy traffic, while effort is the larger lever. To save $1,000 a month you need either about 9,300 cacheable tokens or a single effort step.

If your bill is too high, the larger lever is almost always effort, not caching. On the reference shape — 4,000 in, 900 out, Claude Sonnet 5.5 — caching the stable prefix saves roughly $324 a month at high, while dropping from high to low saves roughly $3,780. But caching costs nothing in quality, so the order is not a choice: cache first, then move effort.

The two levers move different sides of the bill, which is why the comparison is shape-dependent rather than absolute.

What each lever actually moves

Caching moves the input side. Effort moves the output side. On the reference shape the input side is only 10% of the bill at high, so caching has a low ceiling there. Flip the shape to a long-context retrieval task and the ranking reverses, because the input portion becomes most of the bill.

leverside of the bill it moveswhat changesceiling
prompt cachinginput$2.00 down to $0.20 per million on cached tokensthe whole input portion
lower effortoutput1x to 45x modelled thinking tokensmost of the output portion

The ceiling is the thing to estimate before you optimise. If input is 10% of your bill, no amount of caching saves more than 10%. If output is 90%, effort is where the money is.

What prompt caching saves

On Claude Sonnet 5.5 the three relevant rates are $2 per million for fresh input, $0.20 per million for a cache read, and $2.50 per million for a cache write. A read is 90% cheaper than fresh input. A write is 25% more expensive than fresh input.

Take the reference shape with a stable 3,000-token prefix reused on every call, and 1,000 tokens of genuinely new input:

That is a saving of $0.0054 per call, or about 67% of the input cost. At 2,000 calls a day over 30 days, 60,000 calls, the monthly saving is about $324.

Notice what caching does not do: it does not change the answer. The model sees the same tokens, rendered the same way, at a lower price. That makes it the safest lever on the list, and the first one to pull.

What lowering effort saves

Now run the same arithmetic on the output side, holding the prompt fixed. From the reference shape at Claude Sonnet 5.5 rates:

effortcost / call60,000 calls / monthsaving vs high
low$0.0170$1,020$3,780
medium$0.0350$2,100$2,700
high$0.0800$4,800baseline
xhigh$0.1880$11,280minus $6,480
max$0.4130$24,780minus $19,980

Dropping from high to low saves $3,780 a month on the same 60,000 calls. That is about 11.7x the caching saving on the identical shape, and it comes from one string in the request body.

The catch is real, though. Caching is free quality-wise and effort is not. Lowering effort changes the answer, so it has to be validated against your own evals before it goes to production. That is the whole trade: a larger saving for a cost that is not measured in dollars.

Saving $1,000 a month: two routes

Suppose you need to take $1,000 a month out of this workload. Over 60,000 calls that is $0.01667 per call. Here is what each route has to do to get there.

routewhat it takesquality cost
cache the stable prefixabout 9,300 tokens of reusable prefix at a 90% hit ratenone, the output is identical
drop effort from high to medium$0.045 saved per call, clearing the target 2.7x overthe answer changes
drop effort from high to low$0.063 saved per call, clearing the target 3.8x overthe answer changes more

The cache figure comes from the read discount: each cached token saves $1.80 per million, so $1,000 over 60,000 calls needs about 9,300 cached tokens. That is a large prefix. A single effort step does the same job with no prompt engineering at all, provided the quality holds.

If you have both a large prefix and slack in the effort level, you do not have to choose. Take the caching saving, then take the effort saving on top.

Cache write is a premium, not a discount

It is easy to treat caching as uniformly cheaper. It is not. The write costs more than fresh input.

Caching therefore pays from the second reuse onward. If the prefix changes on every request, every call is a write and you pay 25% more than you would have paid without caching at all. A prefix that is rebuilt per request is the one case where caching is a straight loss.

Where the two levers interfere

The interference is a single mechanism: effort shapes the rendered prompt. Change it at the top level and the cached prefix no longer matches, so the next request is a cache miss and you pay a write again.

If you vary effort per request at the top level, you turn a workload with a 90% hit rate into a workload with a 0% hit rate. You then pay the write premium on every call, and the caching lever does nothing.

The fix is per-message effort. Setting effort on a message rather than on the request keeps everything before that message identical, so the prefix stays cached and only the new work is affected.

Do not combine top-level effort changes with prompt caching. Either fix the effort level for the request, or move the variation down to per-message effort.

When caching is the bigger lever

The ranking reverses on input-heavy traffic, which is why the shape matters more than the lever. Take a retrieval task with 100,000 input tokens and a 500-token answer. Input at $2 per million is $0.20 per call. Output at high, with the modelled 8x multiplier, is 4,000 tokens at $10 per million, or $0.04. Input is 83% of the bill, and effort can only reach the other 17%.

leversaving per callwhat it touches
cache a 90,000-token prefix$0.162the input side
drop high to low$0.035the output side

Caching the 90,000-token prefix cuts it from $0.18 to $0.018 and leaves $0.02 of fresh input, so the input side falls from $0.20 to $0.038. The effort lever, on the same call, can only touch the $0.04 output portion. Caching is roughly 4.6x the saving here, the mirror image of the reference shape.

The rule that falls out is simple: look at the input share of your bill before you decide which lever to pull. Above roughly 50%, cache. Below it, move effort.

The decision rule

  1. Cache any prefix that is reused two or more times. It is free quality-wise, so there is no reason not to.
  2. Do not cache a prefix that changes on every request. You would be paying the write premium for nothing.
  3. Once caching is in place, move effort. That is where the large multiples live.
  4. Never vary top-level effort on a cached prefix. Use per-message effort if you need to vary it at all.
  5. Re-check the ranking whenever the shape changes. Input-heavy traffic and output-heavy traffic have opposite answers.

Model both levers side by side in the calculator: it prints the per-call and monthly cost for every effort level, and the input share tells you which lever has room to move.

The short version

Frequently asked

Which saves more, prompt caching or lowering the effort level?

On output-heavy traffic, lowering effort by far. On the reference shape, dropping from high to low saves about $3,780 a month while caching the stable prefix saves about $324. Caching wins only when the input side dominates the bill.

Is caching worth it if the prefix is only reused twice?

Yes, and barely. Two uncached calls cost $4 per million tokens of prefix, while a cache write plus a read costs $2.70, so caching pays from the second reuse onward. Anything less than a second reuse is a loss.

Why does my cache stop working when I change the effort level?

Effort shapes the rendered prompt, so changing it at the top level invalidates the cached prefix. Use per-message effort instead, which leaves everything before that message cached.

Can I use both levers at once?

Yes, and you should. Cache the stable prefix first because it costs nothing in quality, then move effort on the traffic that does not need the higher level. Just never vary top-level effort on a cached prefix.