Artificial Analysis’s first independent evaluation confirms Fable 5.1 is the most capable model measured — and updates yesterday’s cost framing: the 75% cache-read cut saves ~$1.40 per task, not enough to offset ~1.7× the output tokens.
What Happened
Artificial Analysis published its first independent evaluation of Claude Fable 5.1 — and it does two things at once. On capability, the result is unambiguous: Fable 5.1, run at maximum effort with Anthropic’s default server-side fallback enabled, scores 66 on the firm’s Intelligence Index — the highest AA has ever recorded. That puts it ahead of Claude Opus 5 at 63, the prior Fable 5 at 62, and both GPT-5.6 ‘Sol’ and Grok 4.6 at 61. All figures are Artificial Analysis’s own, on its own harness and methodology.
On cost, the picture is more complicated — and it directly updates the framing this outlet led with yesterday. By AA’s accounting, Fable 5.1 costs approximately $3.7 per Intelligence-Index task — the firm’s article states $3.76 while its live model page shows $3.69, a discrepancy worth flagging rather than resolving, likely a post-publication recompute. Either way, the right figure to work with is the range: roughly $3.7. That is about 20% more than Fable 5’s ~$3.1 per task, and approximately 1.6 times Opus 5’s $2.34. The driver, per AA: Fable 5.1 uses roughly 1.7 times the output tokens of its predecessor, and output is billed at $50 per million tokens — the most expensive line in Anthropic’s pricing stack.
The 75% cut to cache-read pricing ($1.00 → $0.25 per million tokens) — Anthropic’s headline cost move at launch — saves approximately $1.40 per task on AA’s workload. That saving is real, but it cannot offset the extra output spend. The result: capability went up, and on this specific workload, cost went up too.
The key insight: Anthropic’s ‘cheaper’ claim and Artificial Analysis’s ‘+20% per task’ finding are not contradictions — they measure different things. Anthropic’s saving is cache-read-conditional: it dominates on long-running agents re-reading large static contexts. AA’s cost rose because Fable 5.1 generates ~1.7× the output tokens, and output at $50/MTok dwarfs what the cache cut saves. The sticker story and the workload story diverge. That divergence is the thing to internalize.
The Structural Read
Yesterday’s read on Fable 5.1 — that Anthropic’s cache-read cut was a targeted reprice of agentic workloads — was correct about the mechanism. The error was in the inference: that the model would therefore be cheaper to run. Artificial Analysis’s first independent per-task measurement shows capability and cost rising together. The efficiency narrative needs that asterisk, and this outlet is adding it now.
Three structural points follow from the data.
First: cache-conditional ‘cheaper’ is a real category, not a marketing hedge. Anthropic’s saving is genuine — but it accrues specifically where cache reads are the dominant cost driver: long-horizon agents that repeatedly read a large, stable context. On AA’s benchmark workload, that condition does not hold cleanly, because Fable 5.1 also thinks and writes substantially more. Output at $50/MTok is the most expensive component. When output volume rises ~1.7×, a 75% cut to the cheaper cache-read line saves ~$1.40 per task — not enough. The lesson is not that Anthropic’s claim was false; it is that ‘cheaper’ is always workload-conditional, and the sticker story and the actual-bill story can diverge sharply depending on which line items grow.
Second: capability and cost both rose, and that is the honest frontier story right now. A more capable model costs more to run on a per-task basis — not because pricing went up, but because a more capable model uses more compute to produce more output. The efficiency gains are real at the cache layer; they are not yet sufficient to offset the cost of doing more work. Labs will continue to claim ‘cheaper’ at launch because the cache story is true for their target customer. Independent per-task measurement is how the rest of the market checks whether that holds for their workload.
Third — and this is the subtler point — the number-one benchmark score is not a clean single-model result. Artificial Analysis ran Fable 5.1 with Anthropic’s default server-side fallback active. Approximately 4% of the output tokens in the benchmark were served by Opus 4.8 or Opus 5, routing safety-flagged requests away from Fable 5.1 itself. The 66 on the Intelligence Index is therefore the score of Fable 5.1 plus Anthropic’s routing system — what this outlet frames as a blended-system score, not AA’s terminology. AA does not quantify what the score would be without the fallback, and this piece will not speculate on that number either.
But the implication is durable: as labs ship models wrapped in routers, safety fallbacks, and effort-level settings, a benchmark increasingly measures a serving stack, not a model. ‘Model X is number one’ is quietly becoming ‘model X’s product configuration is number one.’ That is, in many ways, fairer to how the model is actually used in production. It is also less clean than the single-number headline suggests, and anyone buying on benchmarks should internalize it.
Map of AI — Benchmarks Measure Systems, Not Models
The leaderboard is testing the product
When a lab ships a model with a default router, effort settings, and a safety fallback — and that is what Anthropic ships by default — the benchmark score is a system score. The ~4% of Fable 5.1’s AA benchmark tokens that came from Opus 4.8/Opus 5 are a small number with a large structural implication: the frontier competition is increasingly between serving stacks, not weights. The Map of AI framework tracks this shift at the infrastructure and orchestration layers, where routing decisions are made and where cost is ultimately set.
Benchmark Component Scores (AA Harness, AA Configuration — Not Comparable to Anthropic’s Own Figures)
Note: These scores are on Artificial Analysis’s own harness and configurations. They are not directly comparable to Anthropic’s launch-post figures — which used different benchmarks and settings (e.g. Anthropic reported HLE 60.9–65%, Terminal-Bench 4.0 55.8%). Label the harness or the comparison is meaningless.
Three Implications
IMPLICATION 1 — FOR BUYERS: Always specify your workload before the benchmark
Fable 5.1 at ~$3.7/task on AA’s Intelligence-Index workload vs. Opus 5 at $2.34 is a real gap. Whether that gap shrinks or widens on your specific workload depends entirely on how much of your bill is cache reads versus output generation. Anthropic’s cache-read saving is substantial for agents that repeatedly read large static contexts; it is less impactful on workloads where output volume is the dominant cost driver. Map your own token distribution before using any launch-day pricing claim as a purchasing signal.
IMPLICATION 2 — FOR EVALUATORS: The leaderboard score is a system score now
The ~4% of Fable 5.1’s benchmark tokens that came from Anthropic’s safety fallback (Opus 4.8 / Opus 5) make the number-one result a blended-system score
91,000+ executives read Business Engineer for the AI strategy frameworks cited by ChatGPT, Claude, and Perplexity.
Sources: artificialanalysis.ai · artificialanalysis.ai · fourweekmba.com · x.com









