Sonnet 5.5 Tops Opus 5.5 on 1 of 8 Anthropic Benchmarks

Every benchmark and price here is Anthropic’s own, from its Sonnet 5.5 and Opus 5.5 pages. The Moonshots clip was transcribed by this publication with automatic speech recognition. This publication verified none of it independently.

A Moonshots podcast clip says Sonnet 5.5 beats Opus 5.5 on Terminal-Bench 4.0, 70.6% to 66.4%, at half the price. Anthropic’s own pages confirm those figures, and also show Opus 5.5 ahead on the other seven rows of the same table, with cache-read prices identical. All figures are Anthropic’s own, the clip was transcribed here by automatic speech recognition, and this publication verified none of it independently.

What the Clip Says

A speaker on the Moonshots podcast, in episode 299, says Sonnet 5.5 “jumped from 10% to 70%” on Terminal-Bench 4.0 “in a single generation,” and that it beats “Anthropic’s top model Opus 5.5 at 66.4% for half the price.” He adds that it is the first Sonnet launched with cyber safeguards, a phrase taken from the clip’s published transcript, because this publication’s own recognition caught only the start of that sentence.

This publication transcribed a clip of the passage with automatic speech recognition, so the wording may contain errors. It checked each figure against Anthropic’s own pages for Sonnet 5.5 and Opus 5.5.

Five of the eight rows in Anthropic's Sonnet 5.5 table: Sonnet 5.5 is ahead only on Terminal-Bench 4.0. Anthro
Five of the eight rows in Anthropic’s Sonnet 5.5 table: Sonnet 5.5 is ahead only on Terminal-Bench 4.0. Anthropic’s own figures; the full table has three more rows where Opus 5.5 also scores higher.

What Anthropic’s Table Says on Terminal-Bench

Anthropic’s Sonnet 5.5 page lists Terminal-Bench 4.0 at 70.6% for Sonnet 5.5, 10.3% for Sonnet 5 and 66.4% for Opus 5.5. The clip’s figures match. This publication’s own arithmetic: the gap between Sonnet 5.5 and Opus 5.5 is 4.2 points, and the rise from Sonnet 5 is 60.3 points.

Anthropic’s Opus 5.5 page says its Terminal-Bench 4.0 result is reported at xhigh effort, with a standard error of plus or minus 2.6 points for Opus 5.5 and 1.6 to 2 points for the other Claude models.

The Other Rows

The same table has seven more rows, and on all seven Opus 5.5 scores higher. FrontierCode 1.1 reads 46.2% at Max effort and 52.1% at Xhigh for Sonnet 5.5, against 54.4% for Opus 5.5. CursorBench 4.0 reads 55.5% against 57.8%, and Humanity’s Last Exam with tools 64.5% against 67.7%.

OSWorld 2.1 reads 80.1% against 81.8%, GDPval-AA 1844 against 1846, AA-Briefcase 1811 against 1822, and Chartography without tools 61.6% against 64.4%. This publication’s own arithmetic: Sonnet 5.5 is ahead on one of the eight rows and behind on seven.

Anthropic writes that Opus 5.5 “remains clearly stronger at complex, open-ended work requiring sustained judgment.” It also says that on CursorBench Sonnet 5.5’s best score is within about two points of Opus 5.5.

The Half-Price Claim

Anthropic prices Sonnet 5.5 at $2 per million input tokens, $10 per million output tokens and $0.20 per million tokens for cache reads. It prices Opus 5.5 at $4 and $20 per million input and output tokens and $0.20 per million for cache reads.

This publication’s own arithmetic: the input and output prices are half, and the cache-read price is the same. Anthropic frames cost per task, not per token, saying Sonnet 5.5 costs up to 30% less per task than Sonnet 5, and that it complements Opus 5.5 best at lower effort settings, where it costs less per task.

Cyber Safeguards and “Top Model”

Anthropic says that because Sonnet 5.5’s cybersecurity capabilities are comparable to Opus 5’s, it is “the first Sonnet model to launch with cyber safeguards.” It says higher-risk cybersecurity tasks will visibly fall back to Sonnet 5.

On “top model,” Anthropic’s news listing says Opus 5.5 performs at the level of Claude Fable 5.1 on most work, and its Opus 5.5 page compares it with Claude Mythos 5.1 on biology and cybersecurity. This publication draws no ranking from that.

What Is Not Established

This publication did not read the effort setting behind Sonnet 5.5’s 70.6%, did not retrieve the table’s footnote text on the Sonnet page, and did not check any independent leaderboard. Every benchmark and price is Anthropic’s own.

The clip is accurate on the figures it quotes and silent on the other seven rows. Nothing here says which model to choose, and nothing in this piece is investment advice.

This piece draws on Anthropic’s Sonnet 5.5 and Opus 5.5 pages and news listing, and on a Moonshots podcast clip that this publication transcribed with automatic speech recognition. Every benchmark and price is Anthropic’s own, and none was independently verified. The point gaps, the one-of-eight count and the price relations are this publication’s own arithmetic. Nothing above predicts anything, and nothing here is investment advice.

Sources: anthropic.com · anthropic.com · Anthropic news listing (anthropic.com · Moonshots podcast episode 299, YouTube Blyb1D927pM (clip 770), transcribed by the publication with macOS ASR

Scroll to Top

Discover more from FourWeekMBA

Subscribe now to keep reading and get access to the full archive.

Continue reading

FourWeekMBA