Anthropic Claude Sonnet 5.5 Ships a Safety Router, Not a Refusal

The second model in Anthropic’s Claude 5.5 family doesn’t just add capability — it routes capability, and that is a structurally different product decision.

What Happened

Anthropic published Claude Sonnet 5.5 on September 28, 2026 — API id claude-sonnet-5-5 — describing it as the second model in the Claude 5.5 family and making it available across all platforms, including Amazon Web Services, Google Cloud, and Microsoft Azure. The list price is $2 per million input tokens and $10 per million output, identical to Sonnet 5. This is not a price cut. What the page claims is “up to 30% less per task” — a result attributed to the model needing far fewer tokens to complete the same work, stated as a finding from Anthropic’s own testing. Both qualifiers — “up to” and “in our testing” — are load-bearing and are kept here.

On Anthropic’s own published benchmarks — all self-reported and not independently verified here — the largest gains are concentrated in agentic, computer-use, and visual work. Terminal-Bench 4.0 moves from 10.3% to 70.6%, a difference of +60.3 percentage points by this publication’s subtraction. Chartography (visual chart recognition without tools) moves from 15.6% to 61.6% (+46.0 pp). OSWorld 2.1, marked partial on the page, goes from 57.0% to 80.1% (+23.1 pp). CursorBench 4.0 moves from 34.1% to 55.5% (+21.4 pp). Humanity’s Last Exam with tools moves from 54.9% to 64.5% (+9.6 pp). GDPval-AA v2.1 and AA-Briefcase v1.1 are index scores — 1449 to 1844 and 1359 to 1811 respectively — and are not mixed into the percentage-point comparisons above. FrontierCode 1.1 Main is excluded from comparison entirely, and the reason is worth spelling out: the page reports Sonnet 5.5 at 46.2% on Max and 52.1% on Xhigh against Sonnet 5 at 42.4%, footnotes effort levels for some entries but not others, and a difference taken across mismatched settings would not mean anything.

One published oddity is worth recording without resolving. On Terminal-Bench 4.0, Sonnet 5.5’s self-reported 70.6% sits above the 66.4% published for Opus 5.5, the more expensive flagship in the same family. Both numbers appear on Anthropic’s page. This piece declines to rank the two models on that benchmark: the page footnotes effort levels for some entries and not others, and nothing available here confirms the two scores were produced at the same setting. Across the rest of the published table Sonnet 5.5 generally sits just below Opus 5.5 — 55.5% against 57.8% on CursorBench, 80.1% against 81.8% on OSWorld, 61.6% against 64.4% on Chartography — which is the ordering a mid-tier model would be expected to show, and it makes the Terminal-Bench line the one that stands out rather than the rule.

Benchmark Movement — Sonnet 5 → Sonnet 5.5 (self-reported, Anthropic)

Terminal-Bench 4.0 10.3% → 70.6%
Chartography (no tools) 15.6% → 61.6%
OSWorld 2.1 (partial) 57.0% → 80.1%
CursorBench 4.0 34.1% → 55.5%
Humanity’s Last Exam (with tools) 54.9% → 64.5%

Point differences are this publication’s subtraction from Anthropic’s published figures. FrontierCode 1.1 excluded: effort-level footnotes are inconsistent across model rows, making comparison unreliable.

The key insight: Anthropic’s two new safety controls in Sonnet 5.5 are not about what the model can do. They are both about what comes out of it — which capability a customer can reach, and which reasoning a customer can carry away. That is a different kind of safety architecture than the industry has shipped before.

The movement is concentrated in agentic, computer-use and visual work. On knowledge work with tools the same p
The movement is concentrated in agentic, computer-use and visual work. On knowledge work with tools the same pair of models is under ten points apart.

The Structural Read

The line that deserves the most attention in Anthropic’s Sonnet 5.5 announcement is not a benchmark number. It is this: “Users can still find and fix bugs in their code as part of routine software development, but higher-risk cybersecurity tasks will visibly fall back to Sonnet 5.” Most safety machinery in this industry is binary — the model answers, or it declines. This one degrades. The request is still served, by the older and less capable model, and the user is told it happened.

That makes capability a graded, routable property rather than a fixed one. The product stops being one model with a policy attached and becomes a ladder of models with a router, where the policy chooses the rung. One commercial consequence follows directly from that design: an older model stops being purely legacy and becomes part of the safety architecture of the newer one. Sonnet 5 is not retired by Sonnet 5.5. It is enlisted by it.

It is important to be precise about what is not known here. The page does not define what counts as “higher-risk.” It does not say who or what makes that determination. It does not describe how the fallback is surfaced to the user beyond the word “visibly.” None of that is supplied in this piece, because the page does not supply it.

Anthropic — Claude Sonnet 5.5 Product Page

“Sonnet 5.5 is the first Sonnet model to launch with safety classifiers that prevent reasoning extraction… Claude’s thinking cannot be decoupled from the account that created it.”

The second new control points in a different direction, and the two together are the actual structural news. A classifier whose job is to stop reasoning from being extracted is not protecting the person asking the question. It is protecting the trace. “Expanded preserved thinking” means the chain-of-thought is bound to the originating account and cannot travel independently of it. Read plainly: the reasoning is a proprietary asset, and the classifier is the lock.

Set the two controls beside each other and the pattern is legible. One governs which capability a customer can reach. The other governs which reasoning a customer can carry away. Neither is about what the model can do. Both are about what comes out of it. No company is named in connection with distillation and nothing here alleges that anyone has distilled anything. The biology safeguards, separately, are carried over unchanged from Sonnet 5 — the new controls are the cyber fallback and the reasoning classifiers only.

Product Overhang Doctrine

Capability as a Routable Property

The Product Overhang Doctrine holds that capability accumulates invisibly until it surfaces all at once. Sonnet 5.5 surfaces two things simultaneously: a large capability jump in agentic tasks, and a set of controls that determine how that capability is distributed. The jump and the lock ship together. That is not coincidence — it is product architecture.

On pricing, the quieter change is worth naming precisely. The rate card is fixed — $2 input, $10 output, $0.20 cache reads, $2.50 cache writes per million tokens, unchanged from Sonnet 5. Opus 5.5 sits at $4, $20, $0.20, and $5 respectively. Anthropic also states that Sonnet 5.5 “generates outputs 30%+ faster than Sonnet 5”, making it the company’s fastest Sonnet model to date. On cost, what Anthropic claims is “up to 30% less per task,” attributed to the model requiring far fewer tokens to complete the same work, described as a result from Anthropic’s own testing. When the rate card is fixed and the saving comes from token efficiency, the vendor is competing on cost per task while still billing per token. Those two units pull in opposite directions, and only the second one appears on the invoice. That is a property of the pricing structure, not a criticism of it, and no customer’s actual bill is estimated here.

The product page describes Claude Haiku 5.5 as joining “the Claude 5.5 family in the coming weeks”. It is not shipping today, and nothing here predicts when it does.

Three Implications

IMPLICATION 1 — THE LEGACY MODEL IS NOW INFRASTRUCTURE

When a newer model routes high-risk requests to its predecessor, the predecessor cannot be sunset without disrupting the safety architecture of the successor. Sonnet 5 has a structural role in Sonnet 5.5’s deployment. That changes the economics of model retirement and raises the switching cost for enterprise customers who are effectively relying on two models, not one.

91,000+ executives read Business Engineer for the AI strategy frameworks cited by ChatGPT, Claude, and Perplexity.

Every benchmark figure above is published by Anthropic on its own product page and is self-reported. None of it has been independently replicated here. The list price is unchanged from Sonnet 5, so nothing above is a price cut or a reduction in token rates — the claimed saving is an up-to figure for cost per task, from the company’s own testing, attributed to the model needing fewer tokens for the same work. The page does not define what counts as a higher-risk cybersecurity task, does not say who or what makes that determination, and does not describe how the fallback to Sonnet 5 is surfaced to a user. None of those gaps is filled in above. The percentage-point differences between the two models are this publication’s own subtraction, not figures Anthropic published. FrontierCode is excluded from that comparison because the page footnotes effort levels for some entries and not others, and a difference across mismatched settings would carry no meaning. GDPval-AA and AA-Briefcase are index scores rather than percentages and are not mixed into it. Where Sonnet 5.5’s Terminal-Bench score sits above the figure published for Opus 5.5, both numbers are reported and neither model is ranked above the other, because nothing available here confirms the two were produced at the same effort setting. Claude Haiku 5.5 is described as joining the family in the coming weeks and is not shipping today, and no arrival date is predicted. The biology safeguards are carried over unchanged from Sonnet 5. Nothing above names any party in connection with reasoning extraction or distillation, and nothing above alleges that anyone has distilled any model. Independent replication, parameter counts, architecture, training details, context window, revenue, adoption figures, and any Sonnet 5 retirement date are not established and do not appear — a limit of this reporting rather than evidence that none exist. Nothing above predicts Anthropic, model pricing, any competitor, Haiku, or what any developer does next.

Sources: anthropic.com

Scroll to Top

Discover more from FourWeekMBA

Subscribe now to keep reading and get access to the full archive.

Continue reading

FourWeekMBA