Anthropic published two different savings figures for Claude Opus 5.5 and both are correct — the gap between them is a precise map of where AI cost actually lives.
What Happened
In new documentation published September 22, 2026, Anthropic launched Claude Opus 5.5 with a posted per-token rate cut of 20%: input tokens move from $5 to $4 per million, output tokens from $25 to $20 per million. Separately, the company states the model “will cost 40% less than Opus 5 on typical workloads.” These are not competing claims and neither is an error. They measure fundamentally different things, and the number that connects them is a third figure most coverage will underweight: cache reads are now priced at $0.20 per million tokens, down 60% from Opus 5. Cache reads are the dominant cost driver in agentic and coding workloads — which is exactly the segment Opus 5.5 is positioned for.
Anthropic also reports output generation more than 30% faster than Opus 5. On benchmark performance, the company published its own table — not independently verified — showing Opus 5.5 at 66.4% on Terminal-Bench 4.0 (vs. 52.3% for Opus 5), 57.8% on CursorBench 4.0 (vs. 46.6% for Opus 5), 1846 Elo on GDPval-AA v2.1 (vs. 1708 for Opus 5 and 1735 for Fable 5.1), and 81.8% on OSWorld 2.0 (vs. 74.0% for Opus 5 and 80.7% for Fable 5.1, with both figures marked partial on the source page). Critically, Anthropic itself states that Opus 5.5 “performs at the level of Claude Fable 5.1 on most work” and explicitly notes that benchmark margins are unreliable at these capability levels.
That last sentence deserves to travel as far as the numbers. The vendor publishing the favourable table is simultaneously telling readers the gaps in that table are not reliable. That is not a disclaimer to skip past. It is the most structurally honest line in the release, and it changes how every figure on the page should be read.
The key insight: The 40% typical-workload estimate and the 20% per-token rate cut are both true, and the distance between them is exactly explained by the 60% cache-read price cut. Your actual saving lives in your own usage mix — and only your own invoice can tell you which number applies to you.

The Structural Read
A posted rate is a price per unit. It applies identically to every buyer, and it is the number that belongs in a price index. A typical-workload estimate is a weighted average across an assumed usage mix — which means its denominator is not tokens, it is a mix of tokens whose proportions belong to Anthropic’s assumed buyer, not yours. These are different objects. Treating them as competing claims about the same thing is the error; recognizing that the distance between them is checkable arithmetic is the correction.
The checkable arithmetic: cache reads fell 60% while headline rates fell 20%. Cache reads are the bulk of what agentic and coding work actually consumes — long context windows, repeated system prompts, tool call results flowing back through the context repeatedly. If your workload is structured that way, the 60% cut on the thing you mostly consume produces a blended saving that trends toward the 40% estimate. If your workload is not structured that way — if you are making short, stateless calls with minimal context reuse — the 60% cache cut is nearly irrelevant and your saving is closer to the 20% rate cut. Both buyers can truthfully say Anthropic’s stated figures apply to them. Neither can truthfully say the other figure applies to them. The mix is the whole question, and the mix is yours to know, not Anthropic’s to publish.
There is a second structural issue that the benchmark table surfaces. Anthropic’s own figures — vendor-reported, not independently verified — show meaningful gains on Terminal-Bench 4.0, CursorBench 4.0, GDPval-AA v2.1, and OSWorld 2.0 (the latter marked partial on the source page). The company then says the margins are unreliable at these capability levels. That combination is the honest state of frontier evaluation in 2026: something improved, the direction is real, and the size of any gap relative to a competitor is a number no one — including the people who measured it — should be treating as durable. Level-based comparisons age at the rate the leading model improves, and that rate is fast.
Product Overhang Doctrine — Applied
“When both numerator and denominator move at once — posted price falls, reported capability rises — a price index finally registers something. That is an improvement on registering nothing. But it still cannot separate how much of the gain came from price and how much from capability. The figure most likely to travel through coverage is the mix-dependent one. The measurement problem does not disappear when the price moves. It just becomes harder to notice, because now there is a number that looks like it answers the question.”
This is the Product Overhang Doctrine in its pricing variant. Capability has been accumulating — Opus 5, Fable 5.1, now Opus 5.5 — and each release surfaces that accumulation in a different way. Sometimes through a benchmark jump. Sometimes through a price list that finally moves. Sometimes, as here, through both simultaneously. The analytical trap is assuming that because a number moved, the picture is now legible. It is more legible. It is not complete. The cache-read architecture is the hidden variable that makes the 40% figure coherent, and it was not visible in the headline price at all.
Three Implications
IMPLICATION 1 — FOR BUYERS AND PROCUREMENT
Neither the 20% rate cut nor the 40% workload estimate tells you your saving. Your saving is a function of your cache-read share. If you are running agentic pipelines, coding agents, or any workflow with high context reuse, your blended reduction will trend toward the larger figure. If you are not, it will not. The only way to know is to audit your own token-type distribution before any budget conversation — not after the invoice arrives.
IMPLICATION 2 — FOR READING BENCHMARK TABLES
All benchmark figures cited here are vendor-reported on Anthropic’s own table, not independently verified. Anthropic itself states that benchmark margins are unreliable at these capability levels — which means the directional signal (something improved) is informative, but the specific margin against any competitor is not a number to spend on. OSWorld 2.0 figures are additionally marked partial on the source page. Treat direction as real; treat magnitude as noise until independent evaluation catches up.
IMPLICATION 3 — FOR COMPETITIVE ANALYSIS
Any level-based comparison ages at the rate the leading model improves — and that rate is fast. Anthropic’s own table shows Terminal-Bench 4.0 moving from 52.3% to 66.4% between its own two model generations. Whatever those margins are worth given the company’s own reliability caveat, the general property is structural: a competitive ranking published today describes a world that may not exist by the time the team reading it acts on it. Ranking-as-snapshot is the wrong frame. Rate-of-improvement is the right one.
Editorial Transparency
On the benchmarks in this article
All benchmark figures (Terminal-Bench 4.0, CursorBench 4.0, GDPval-AA v2.1, OSWorld 2.0) are drawn from Anthropic’s own published table — they are vendor-reported, not independently verified. Anthropic explicitly states that benchmark margins are unreliable at these capability levels. OSWorld 2.0 figures are marked partial on the Anthropic source page. Nothing in this article declares any model ahead of, or better than, any other. This article is not investment advice.
91,000+ executives read Business Engineer for the AI strategy frameworks cited by ChatGPT, Claude, and Perplexity.
All benchmark figures above are Anthropic’s own reported numbers, not independent evaluations, and Anthropic itself notes that benchmark margins are unreliable at these capability levels — so the direction is informative and the margins are not. This is not investment advice. The OSWorld 2.0 figures are marked partial on the source page. The 20% figure is the posted per-token rate cut; the 40% figure is Anthropic’s estimate for typical workloads at default settings and is not a price cut — it is larger because cache reads fell 60% and cache reads dominate agentic and coding spend, so any given buyer’s saving depends on their own usage mix, which neither published number reveals. Nothing above declares any model best or ahead of any other, makes any claim about model routing or which model serves a given request, or cites any independent benchmark, latency, throughput, revenue, cost-to-serve or adoption figure. Nothing is predicted.
Sources: anthropic.com









