Cost

Seven levers that cut a Claude API bill, ranked by how much they save

Matching effort to traffic class is the largest lever on output-heavy work, caching is the safest, and the rest range from a few hundred to a few thousand dollars a month.

Matching the effort level to the traffic class is the largest lever, and it is usually the first one worth pulling. On the reference shape — 4,000 in, 900 out, Claude Sonnet 5.5 — dropping from high to low takes the bill from about $4,800 a month to about $1,020, a 79% cut. Caching the stable prefix is the safest lever and saves about $324 on the same shape. The remaining five range from a few hundred to a few thousand dollars a month depending on your traffic.

The list below is ordered by how broadly each lever applies rather than by raw size on every shape, because the size depends on whether your traffic is input-heavy or output-heavy. The table shows the size on the reference shape so the comparison is concrete.

How the ranking was built

Every figure assumes the reference shape: 4,000 input tokens, 900 output tokens, 2,000 calls a day, 60,000 calls a month, Claude Sonnet 5.5 at $2 per million in and $10 per million out. Change the shape and the ranking changes with it.

rankleversaving on the reference shapebest whenavoid when
1Match effort to traffic class$3,780 / mo (79%)output-heavy traffic running at the defaultthe task genuinely needs the reasoning
2Drop classification and extraction to low25% on a short-output classifierthe task is mechanicaloutputs are long and need reasoning
3Cache the stable system prompt$324 / mo on this shapethe prefix is reused two or more timesthe prefix changes every call
4Move deterministic work to a cheaper model$4,800 / mo from Opus to Sonnetthe output is mechanicalquality is the product
5Trim context that is not earning its place$120 / mo per 1,000 tokens removedthe prompt carries stale contextthe context changes the answer
6Batch the traffic that is not in a hurryvendor-specific, plus a lower effortnobody is waitinga user is waiting
7Cap max_tokensup to $195 / mo at a 1% runaway rateoutputs sometimes run longlegitimate answers are long

Lever 4 is larger than lever 1 on this shape, and lever 6 is unquantifiable from a price list. Treat the order as a reading list, not a strict ranking.

Lever 1: match effort to traffic class

This is the largest lever because effort scales every output token, and output is most of the bill on output-heavy traffic.

effortcost / call60,000 calls / month
low$0.0170$1,020
medium$0.0350$2,100
high$0.0800$4,800
xhigh$0.1880$11,280
max$0.4130$24,780

The default is high on most models and medium on Opus 5.5, so a lot of traffic is paying for reasoning it never asked for. Setting effort to the model default is identical to omitting the parameter, which is why the change is invisible until the invoice arrives.

Do not use this lever on a task that genuinely needs the reasoning. Splitting a hard planning task into a cheaper level does not save money, it moves the cost to a human who has to fix the output.

Lever 2: drop classification and extraction to low

Classification and extraction are the classic low workloads: one right answer, no reasoning required, and a cheaper model usually gets them right.

The arithmetic is smaller than people expect, though, because short outputs mean the input dominates. A classifier on 2,000 in and 20 out costs $0.0056 at high and $0.0042 at low, a saving of 25% rather than 79%.

At 60,000 calls that is about $84 a month. Worth doing, but the bigger win on a short-output task is the model choice and trimming the input, not the effort level.

Do not use this lever when the extraction is ambiguous or the schema is open-ended. That is reasoning work wearing a classification label.

Lever 3: cache the stable system prompt

Cache reads cost $0.20 per million against $2.00 for fresh input on Sonnet 5.5, a 90% discount. Cache writes cost $2.50, a 25% premium, so caching pays from the second reuse onward.

With a stable 3,000-token prefix, the input cost drops from $0.0080 to $0.0026 per call, a saving of about $324 a month on this shape. On input-heavy traffic the same lever is worth far more, because it is attacking the majority of the bill.

Do not use this lever when the prefix changes on every request. Every call becomes a write, and you pay 25% more than you would have paid without caching. Also remember that changing effort at the top level invalidates the prefix, so use per-message effort if you need both.

Lever 4: move deterministic work to a cheaper model

Model choice is a bigger multiple than effort, because both the input and the output rate change at once.

modelinput / output ratecost / call at high60,000 calls / month
Claude Sonnet 5.5$2 / $10$0.0800$4,800
Claude Opus 5.5$4 / $20$0.1600$9,600
GPT-6 Astra$10 / $50$0.4000$24,000

Moving this workload from Opus 5.5 to Sonnet 5.5 at the same effort saves $4,800 a month, which is larger than the effort lever on the same shape. Note that Sonnet 5 and Sonnet 5.5 are priced identically, so switching between those two is not a cost decision at all.

Do not use this lever when quality is the product. The comparison to run first is a stronger model at moderate effort against a weaker model at high effort, because the former is often cheaper.

Lever 5: trim context that is not earning its place

Input tokens cost $2 per million on Sonnet 5.5, so every 1,000 tokens you remove saves $0.002 per call, or about $120 a month at 60,000 calls.

The saving is proportional and predictable, which makes this the easiest lever to forecast. It is also the one that most often finds free money, because prompts accumulate stale context: old examples, superseded instructions, documents that no longer matter.

Do not use this lever when the removed context changes the answer. Trim, then re-run the eval. A prompt that is 20% cheaper and 5% worse is not an optimisation.

Lever 6: batch the traffic that is not in a hurry

Batching removes the latency constraint. That matters for cost because latency is what forces you into a higher effort level in the first place: if nobody is waiting, you can accept a slower turnaround and a cheaper setting.

The direct batch discount is set by the vendor and is not part of the price table used here, so treat the dollar figure as unknown until you check the vendor's current page. The indirect saving is real and quantifiable: it is whatever the effort step down is worth on the batched traffic.

Do not use this lever on anything a person is watching. Batching a synchronous chat endpoint does not save money, it breaks the product.

Lever 7: cap max_tokens

A cap does not change the average, it changes the worst case. A runaway request at max can generate around 40,500 output tokens, which bills at about $0.405. Capping the response at 8,000 tokens limits the same request to about $0.080.

If 1% of 60,000 calls run away, that is 600 calls and about $195 a month recovered. More importantly, a cap makes the bill predictable, which is what lets you forecast it.

Do not set the cap so tight that it truncates legitimate answers, and remember that thinking counts against it. A cap sized for the answer alone is the reason a high-effort request returns nothing.

Do not roll all seven out at once. Change one lever, measure the bill and the quality, then move to the next. Bundled changes make it impossible to tell which one helped.

The short version

Estimate each lever against your own token counts in the calculator before you change anything.

Frequently asked

What is the single biggest lever on a Claude API bill?

Matching the effort level to the traffic class. On the reference shape, dropping from high to low saves about $3,780 a month out of $4,800, which is 79%. No other single change comes close on output-heavy traffic.

Which lever is the safest?

Prompt caching. It changes the price and not the answer, because the model sees the same tokens rendered the same way. It pays off from the second reuse onward.

Is switching between Sonnet 5 and Sonnet 5.5 a cost lever?

No. Both are priced at $2 per million input and $10 per million output, so the choice between them is a quality and recency decision rather than a cost one.

Do I need to change all seven at once?

No, and you should not. Take the free ones first, caching and the max_tokens cap, then test the effort change against your own evals before rolling it out.