GPT-6 Astra and Claude Fable 5.1 Lead the Epoch Capabilities Index — But the Error Bars Tell a Different Story

When confidence intervals overlap across every panel of a fifty-benchmark composite, a ranking read off the point estimates is carrying more weight than the measurement supports — and the domain split underneath the headline is subject to exactly the same caution.

Epoch Capabilities Index — Point Estimates Read From Published Chart (CC-BY)

General Capability

GPT-6 Astra: ~166

Claude Fable 5.1: ~164

GPT-5.6 Sol: ~162

Kimi K3: ~158

Mathematics

GPT-6 Astra: ~170

Claude Fable 5.1: ~166

GPT-5.6 Sol: ~163

Kimi K3: ~158

Software Engineering

Claude Fable 5.1: ~167

GPT-6 Astra: ~164

Kimi K3: ~162

GPT-5.6 Sol: ~160

Index Construction

50+

Distinct benchmarks composited; harder benchmarks raise scores; 90% confidence intervals published

All figures are point estimates read from Epoch AI’s published chart (CC-BY licence) — not from underlying data files. Treat as approximate, not precise to the unit. Exact confidence interval bounds are not stated; they are read from a graphic.

What Happened

Epoch AI’s Capabilities Index, published under a CC-BY licence, now shows four frontier models in a narrow band across three domain panels. Reading point estimates from the published chart: GPT-6 Astra sits at approximately 166 on general capability, Claude Fable 5.1 at around 164, GPT-5.6 Sol near 162, and Kimi K3 at roughly 158. On mathematics the spread widens only slightly, with Astra reaching approximately 170 and Fable 5.1 around 166. These are estimates read from a graphic, not figures extracted from Epoch’s underlying data files, and they should not be treated as precise to the unit.

The ECI is a composite drawn from more than fifty distinct benchmarks, stitched together by comparing models evaluated across several of them — a construction that means stronger performance on harder benchmarks raises scores. That methodology is Epoch’s disclosed design, and nothing here criticises it. What the methodology cannot do is eliminate uncertainty, and Epoch does not claim otherwise: the published chart carries 90% confidence intervals around every point estimate, and those intervals overlap substantially between the leading models in every panel.

The software engineering panel is where the ordering changes most visibly. Fable 5.1 leads at approximately 167, ahead of Astra at around 164, with Kimi K3 near 162 and Sol near 160 — an inversion of the general capability ranking, and one in which Fable 5.1’s confidence interval is conspicuously wide. That width is not a flaw in Epoch’s reporting; it is honest communication about how much the data actually constrains the estimate.

ECI Context — Authorised Timeline

Prior ECI Cycle

Epoch reports Claude Fable 5 at approximately 161 on the ECI — one point ahead of GPT-5.5 Pro, described as Anthropic’s first ECI lead in over a year. A one-point margin on a fifty-benchmark composite is simultaneously a genuinely interesting datum and a genuinely weak foundation for a narrative about field leadership.

September 2026 — Current Reading

Four models cluster within overlapping confidence intervals across general capability while diverging by domain. On point estimates the domain orderings differ; the intervals overlap there too, so that split is suggestive rather than shown.

The key insight: When the lead sits inside the error bar, the leaderboard is measuring something other than capability. A two-point difference between models whose confidence intervals overlap across most of their range is not a demonstrated ordering — it is the best available point estimate with an honest caveat attached. Epoch attaches that caveat visibly, which is to its credit. The problem is entirely downstream, in how the industry consumes the number.

Four models, three domains, three different orderings. On software engineering the ranking inverts entirely. S
Four models, three domains, three different orderings. On software engineering the ranking inverts entirely. Set against confidence intervals that overlap across most of their range, this is a cluster rather than a leaderboard — which is a different thing to buy from.

The Structural Read

The industry has a reliable habit with benchmark scores: it reads the point estimate and discards the interval. That is not a quirk of individual behaviour; it is what happens structurally when a ranked list is the output that travels. A ranked list has no slot for uncertainty. “GPT-6 Astra leads the ECI” is a sentence. “GPT-6 Astra and Claude Fable 5.1 have overlapping 90% confidence intervals on general capability, making a demonstrated ordering unavailable from this data” is a paragraph that does not survive the journey from chart to press release to procurement memo.

This matters commercially because procurement decisions, marketing claims, and in some cases valuations are made on the ordering rather than the interval. A distinction the data does not support ends up carrying commercial weight it cannot bear. That is not a pedantic complaint about statistical communication. It is a description of how benchmark scores become inputs to business decisions that they were never designed to support.

The domain split is where the analysis wants to go next, and it is worth being disciplined about it, because the same objection applies. On the point estimates, Astra sits ahead on general capability and mathematics, Fable 5.1 sits ahead on software engineering, and the SWE ordering inverts further down with Kimi K3 above Sol. But the intervals overlap in that panel too — Fable 5.1’s is conspicuously wide — so the domain split is suggestive rather than established. That is not a reason to ignore it. It is a reason to treat it as a hypothesis worth testing against your own workload rather than a fact to plan around. If different models lead different domains, then the question “which model is best” has no answer that is useful to anybody. The question that replaces it is “best at what, for whom, and at what price” — and that is a procurement question, not a leaderboard question. Price is not in this data at all, so nothing here ranks models on cost or value for money.

Business Engineer — Cluster as Commodity

“A cluster is a commodity in waiting. When frontier models converge within measurement error on general capability, differentiation has to be found somewhere other than the headline number — in price, latency, tooling, reliability, or distribution. The domain split is the signal that tells buyers where to look instead: route by task, not by brand.”

There is history here worth holding. Epoch previously reported Claude Fable 5 reaching approximately 161 on the index and beating GPT-5.5 Pro by a single point — described at the time as Anthropic’s first ECI lead in over a year. A one-point margin on a fifty-benchmark composite is simultaneously a genuinely interesting datum and a genuinely weak foundation for a narrative about who leads the field. Both of those were true then, and both are true now. What is consistent is not which lab is ahead, but what happens to the uncertainty: the interval gets dropped somewhere between the chart and the conversation about the chart, reliably, in every cycle. The index does not claim to measure usefulness, intelligence, safety, or commercial value. It measures performance across a benchmark suite, and Epoch says so.

Three Implications

RANKING DISCARDS THE UNCERTAINTY THE CHART DREW

Every time a benchmark score travels as a ranking rather than as a point estimate with an interval, commercial decisions inherit false precision. Procurement teams, marketing leads, and analysts consuming these numbers as a clean ordering are working from a signal that the underlying data explicitly does not support. The discipline required is to ask, before acting on a leaderboard position, whether the lead is demonstrated or whether it sits inside the error bar — because in this data, for the leading models, it consistently does the latter.

THE DOMAIN SPLIT MAKES “BEST” ILL-POSED AS A QUESTION

On point estimates Astra sits ahead on general capability and mathematics, Fable 5.1 on software engineering, and Kimi K3 above Sol on SWE — all with overlapping intervals, so none of these orderings is established. Even as unproven orderings, they point the same way: “which model is best” collapses under domain specificity. Enterprise buyers who have already been routing workloads by task rather than by brand are ahead of where the leaderboard conversation has arrived. The ECI domain breakdown makes that routing logic legible in a way the composite headline score does not.

CAPABILITY CLUSTERS SHIFT WHERE DIFFERENTIATION CAN LIVE

When frontier models converge within measurement error on general capability, the headline number loses its power to justify a vendor choice. Differentiation — commercially durable differentiation — has to move somewhere else: price, latency, tooling integration, uptime reliability, or distribution reach. This is consistent with previously reported enterprise buying behaviour, not a prediction derived from this data. Nothing here claims these models are substitutable for any particular use case; that depends entirely on the task and on factors the ECI does not measure, including cost.

Business Engineer Framework

The Map of AI — Where Clusters Form and Where Differentiation Moves

The Map of AI traces nine layers of the AI stack and more than 200 companies sitting across them. When frontier model scores converge within measurement error — as the ECI now shows — the map identifies where competitive separation actually lives: not in the capability layer where the cluster formed, but in the distribution, tooling, and reliability layers below it. The domain split the ECI reveals is exactly the kind of signal the Map of AI is built to translate into a structural position.

91,000+ executives read Business Engineer for the AI strategy frameworks cited by ChatGPT, Claude, and Perplexity.

The index figures in this article are point estimates read from a chart published by Epoch AI under a CC-BY licence. They were not obtained from Epoch’s underlying data files and should not be treated as precise to the unit. Epoch publishes 90% confidence intervals alongside these scores, and those intervals overlap substantially between the leading models in every domain shown. No numeric interval bounds are stated here because they are being read from a graphic. Any difference of a few points between models whose intervals overlap should not be read as a demonstrated ordering. Nothing here criticises Epoch’s methodology. Publishing confidence intervals is the opposite of the problem this article describes, which concerns how rankings are consumed downstream rather than how the index is built. The Epoch Capabilities Index measures performance across a benchmark suite; it does not claim to measure usefulness, intelligence, safety or commercial value, and nothing here suggests it does. Price is not part of this dataset. No claim is made about the cost, value for money or cost per unit of capability of any model. No model is described as best, superior or winning, no prediction is offered about future benchmark results or model releases, and nothing here forecasts whether any laboratory takes or retakes a lead. Nothing here claims that any models are substitutable for any particular purpose, which depends on the task. OpenAI, Anthropic and Moonshot AI are private companies. This is business analysis, not investment advice, no view is expressed on any security, and no recommendation is made.

Sources: epoch.ai · epoch.ai · epoch.ai

Scroll to Top

Discover more from FourWeekMBA

Subscribe now to keep reading and get access to the full archive.

Continue reading

FourWeekMBA