Claude Fable 5.1 Tops Artificial Analysis’s Intelligence Index at 66 — and Costs About 20% More Per Task Than Its Predecessor

Artificial Analysis’s first independent evaluation confirms Fable 5.1 is the most capable model measured — and updates yesterday’s cost framing: the 75% cache-read cut saves ~$1.40 per task, not enough to offset ~1.7× the output tokens.

Artificial Analysis Intelligence Index — ~1 Sep 2026 (AA Harness)

66

Fable 5.1 (max + fallback)
Intelligence Index

~$3.7

Cost per Intelligence-Index task
(AA article: $3.76 / model page: $3.69)

+~20%

Cost vs. Fable 5 (~$3.1)
on AA’s per-task workload

~1.7×

Output tokens vs. Fable 5
(drives the cost delta)

What Happened

Artificial Analysis published its first independent evaluation of Claude Fable 5.1 — and it does two things at once. On capability, the result is unambiguous: Fable 5.1, run at maximum effort with Anthropic’s default server-side fallback enabled, scores 66 on the firm’s Intelligence Index — the highest AA has ever recorded. That puts it ahead of Claude Opus 5 at 63, the prior Fable 5 at 62, and both GPT-5.6 ‘Sol’ and Grok 4.6 at 61. All figures are Artificial Analysis’s own, on its own harness and methodology.

On cost, the picture is more complicated — and it directly updates the framing this outlet led with yesterday. By AA’s accounting, Fable 5.1 costs approximately $3.7 per Intelligence-Index task — the firm’s article states $3.76 while its live model page shows $3.69, a discrepancy worth flagging rather than resolving, likely a post-publication recompute. Either way, the right figure to work with is the range: roughly $3.7. That is about 20% more than Fable 5’s ~$3.1 per task, and approximately 1.6 times Opus 5’s $2.34. The driver, per AA: Fable 5.1 uses roughly 1.7 times the output tokens of its predecessor, and output is billed at $50 per million tokens — the most expensive line in Anthropic’s pricing stack.

The 75% cut to cache-read pricing ($1.00 → $0.25 per million tokens) — Anthropic’s headline cost move at launch — saves approximately $1.40 per task on AA’s workload. That saving is real, but it cannot offset the extra output spend. The result: capability went up, and on this specific workload, cost went up too.

The Cost Story Evolves

Launch Day — Anthropic claim

Cache-read price cut 75% ($1.00 → $0.25/MTok). Anthropic cites ~25% cheaper for typical work, up to 45% cheaper for highly agentic tasks. Headline per-token prices unchanged.

~1 Sep 2026 — Artificial Analysis independent eval

Intelligence Index 66 (highest ever). Cost per task: ~$3.7 (+~20% vs Fable 5). Output tokens ~1.7× predecessor. Cache cut saves ~$1.40/task. ~4% of output tokens served by Opus 4.8/Opus 5 fallback.

The honest reconciliation

Anthropic’s saving is cache-conditional — real where cache reads dominate the bill. AA’s measurement is per-task on a blended workload where output costs dominate. These describe different things; neither is wrong, but they cannot be averaged into a single headline.

The key insight: Anthropic’s ‘cheaper’ claim and Artificial Analysis’s ‘+20% per task’ finding are not contradictions — they measure different things. Anthropic’s saving is cache-read-conditional: it dominates on long-running agents re-reading large static contexts. AA’s cost rose because Fable 5.1 generates ~1.7× the output tokens, and output at $50/MTok dwarfs what the cache cut saves. The sticker story and the workload story diverge. That divergence is the thing to internalize.

The Structural Read

Yesterday’s read on Fable 5.1 — that Anthropic’s cache-read cut was a targeted reprice of agentic workloads — was correct about the mechanism. The error was in the inference: that the model would therefore be cheaper to run. Artificial Analysis’s first independent per-task measurement shows capability and cost rising together. The efficiency narrative needs that asterisk, and this outlet is adding it now.

Three structural points follow from the data.

First: cache-conditional ‘cheaper’ is a real category, not a marketing hedge. Anthropic’s saving is genuine — but it accrues specifically where cache reads are the dominant cost driver: long-horizon agents that repeatedly read a large, stable context. On AA’s benchmark workload, that condition does not hold cleanly, because Fable 5.1 also thinks and writes substantially more. Output at $50/MTok is the most expensive component. When output volume rises ~1.7×, a 75% cut to the cheaper cache-read line saves ~$1.40 per task — not enough. The lesson is not that Anthropic’s claim was false; it is that ‘cheaper’ is always workload-conditional, and the sticker story and the actual-bill story can diverge sharply depending on which line items grow.

Second: capability and cost both rose, and that is the honest frontier story right now. A more capable model costs more to run on a per-task basis — not because pricing went up, but because a more capable model uses more compute to produce more output. The efficiency gains are real at the cache layer; they are not yet sufficient to offset the cost of doing more work. Labs will continue to claim ‘cheaper’ at launch because the cache story is true for their target customer. Independent per-task measurement is how the rest of the market checks whether that holds for their workload.

Third — and this is the subtler point — the number-one benchmark score is not a clean single-model result. Artificial Analysis ran Fable 5.1 with Anthropic’s default server-side fallback active. Approximately 4% of the output tokens in the benchmark were served by Opus 4.8 or Opus 5, routing safety-flagged requests away from Fable 5.1 itself. The 66 on the Intelligence Index is therefore the score of Fable 5.1 plus Anthropic’s routing system — what this outlet frames as a blended-system score, not AA’s terminology. AA does not quantify what the score would be without the fallback, and this piece will not speculate on that number either.

But the implication is durable: as labs ship models wrapped in routers, safety fallbacks, and effort-level settings, a benchmark increasingly measures a serving stack, not a model. ‘Model X is number one’ is quietly becoming ‘model X’s product configuration is number one.’ That is, in many ways, fairer to how the model is actually used in production. It is also less clean than the single-number headline suggests, and anyone buying on benchmarks should internalize it.

Map of AI — Benchmarks Measure Systems, Not Models

The leaderboard is testing the product

When a lab ships a model with a default router, effort settings, and a safety fallback — and that is what Anthropic ships by default — the benchmark score is a system score. The ~4% of Fable 5.1’s AA benchmark tokens that came from Opus 4.8/Opus 5 are a small number with a large structural implication: the frontier competition is increasingly between serving stacks, not weights. The Map of AI framework tracks this shift at the infrastructure and orchestration layers, where routing decisions are made and where cost is ultimately set.

Benchmark Component Scores (AA Harness, AA Configuration — Not Comparable to Anthropic’s Own Figures)

Humanity’s Last Exam (AA harness) 59.1%
Terminal-Bench v2.1 (AA harness) 91.4%
SciCode (AA harness) 62.0%

Note: These scores are on Artificial Analysis’s own harness and configurations. They are not directly comparable to Anthropic’s launch-post figures — which used different benchmarks and settings (e.g. Anthropic reported HLE 60.9–65%, Terminal-Bench 4.0 55.8%). Label the harness or the comparison is meaningless.

Three Implications

IMPLICATION 1 — FOR BUYERS: Always specify your workload before the benchmark

Fable 5.1 at ~$3.7/task on AA’s Intelligence-Index workload vs. Opus 5 at $2.34 is a real gap. Whether that gap shrinks or widens on your specific workload depends entirely on how much of your bill is cache reads versus output generation. Anthropic’s cache-read saving is substantial for agents that repeatedly read large static contexts; it is less impactful on workloads where output volume is the dominant cost driver. Map your own token distribution before using any launch-day pricing claim as a purchasing signal.

IMPLICATION 2 — FOR EVALUATORS: The leaderboard score is a system score now

The ~4% of Fable 5.1’s benchmark tokens that came from Anthropic’s safety fallback (Opus 4.8 / Opus 5) make the number-one result a blended-system score

91,000+ executives read Business Engineer for the AI strategy frameworks cited by ChatGPT, Claude, and Perplexity.

Sources: artificialanalysis.ai · artificialanalysis.ai · fourweekmba.com · x.com

Scroll to Top

Discover more from FourWeekMBA

Subscribe now to keep reading and get access to the full archive.

Continue reading

FourWeekMBA