When confidence intervals overlap across every panel of a fifty-benchmark composite, a ranking read off the point estimates is carrying more weight than the measurement supports — and the domain split underneath the headline is subject to exactly the same caution.
What Happened
Epoch AI’s Capabilities Index, published under a CC-BY licence, now shows four frontier models in a narrow band across three domain panels. Reading point estimates from the published chart: GPT-6 Astra sits at approximately 166 on general capability, Claude Fable 5.1 at around 164, GPT-5.6 Sol near 162, and Kimi K3 at roughly 158. On mathematics the spread widens only slightly, with Astra reaching approximately 170 and Fable 5.1 around 166. These are estimates read from a graphic, not figures extracted from Epoch’s underlying data files, and they should not be treated as precise to the unit.
The ECI is a composite drawn from more than fifty distinct benchmarks, stitched together by comparing models evaluated across several of them — a construction that means stronger performance on harder benchmarks raises scores. That methodology is Epoch’s disclosed design, and nothing here criticises it. What the methodology cannot do is eliminate uncertainty, and Epoch does not claim otherwise: the published chart carries 90% confidence intervals around every point estimate, and those intervals overlap substantially between the leading models in every panel.
The software engineering panel is where the ordering changes most visibly. Fable 5.1 leads at approximately 167, ahead of Astra at around 164, with Kimi K3 near 162 and Sol near 160 — an inversion of the general capability ranking, and one in which Fable 5.1’s confidence interval is conspicuously wide. That width is not a flaw in Epoch’s reporting; it is honest communication about how much the data actually constrains the estimate.
The key insight: When the lead sits inside the error bar, the leaderboard is measuring something other than capability. A two-point difference between models whose confidence intervals overlap across most of their range is not a demonstrated ordering — it is the best available point estimate with an honest caveat attached. Epoch attaches that caveat visibly, which is to its credit. The problem is entirely downstream, in how the industry consumes the number.

The Structural Read
The industry has a reliable habit with benchmark scores: it reads the point estimate and discards the interval. That is not a quirk of individual behaviour; it is what happens structurally when a ranked list is the output that travels. A ranked list has no slot for uncertainty. “GPT-6 Astra leads the ECI” is a sentence. “GPT-6 Astra and Claude Fable 5.1 have overlapping 90% confidence intervals on general capability, making a demonstrated ordering unavailable from this data” is a paragraph that does not survive the journey from chart to press release to procurement memo.
This matters commercially because procurement decisions, marketing claims, and in some cases valuations are made on the ordering rather than the interval. A distinction the data does not support ends up carrying commercial weight it cannot bear. That is not a pedantic complaint about statistical communication. It is a description of how benchmark scores become inputs to business decisions that they were never designed to support.
The domain split is where the analysis wants to go next, and it is worth being disciplined about it, because the same objection applies. On the point estimates, Astra sits ahead on general capability and mathematics, Fable 5.1 sits ahead on software engineering, and the SWE ordering inverts further down with Kimi K3 above Sol. But the intervals overlap in that panel too — Fable 5.1’s is conspicuously wide — so the domain split is suggestive rather than established. That is not a reason to ignore it. It is a reason to treat it as a hypothesis worth testing against your own workload rather than a fact to plan around. If different models lead different domains, then the question “which model is best” has no answer that is useful to anybody. The question that replaces it is “best at what, for whom, and at what price” — and that is a procurement question, not a leaderboard question. Price is not in this data at all, so nothing here ranks models on cost or value for money.
Business Engineer — Cluster as Commodity
“A cluster is a commodity in waiting. When frontier models converge within measurement error on general capability, differentiation has to be found somewhere other than the headline number — in price, latency, tooling, reliability, or distribution. The domain split is the signal that tells buyers where to look instead: route by task, not by brand.”
There is history here worth holding. Epoch previously reported Claude Fable 5 reaching approximately 161 on the index and beating GPT-5.5 Pro by a single point — described at the time as Anthropic’s first ECI lead in over a year. A one-point margin on a fifty-benchmark composite is simultaneously a genuinely interesting datum and a genuinely weak foundation for a narrative about who leads the field. Both of those were true then, and both are true now. What is consistent is not which lab is ahead, but what happens to the uncertainty: the interval gets dropped somewhere between the chart and the conversation about the chart, reliably, in every cycle. The index does not claim to measure usefulness, intelligence, safety, or commercial value. It measures performance across a benchmark suite, and Epoch says so.
Three Implications
RANKING DISCARDS THE UNCERTAINTY THE CHART DREW
Every time a benchmark score travels as a ranking rather than as a point estimate with an interval, commercial decisions inherit false precision. Procurement teams, marketing leads, and analysts consuming these numbers as a clean ordering are working from a signal that the underlying data explicitly does not support. The discipline required is to ask, before acting on a leaderboard position, whether the lead is demonstrated or whether it sits inside the error bar — because in this data, for the leading models, it consistently does the latter.
THE DOMAIN SPLIT MAKES “BEST” ILL-POSED AS A QUESTION
On point estimates Astra sits ahead on general capability and mathematics, Fable 5.1 on software engineering, and Kimi K3 above Sol on SWE — all with overlapping intervals, so none of these orderings is established. Even as unproven orderings, they point the same way: “which model is best” collapses under domain specificity. Enterprise buyers who have already been routing workloads by task rather than by brand are ahead of where the leaderboard conversation has arrived. The ECI domain breakdown makes that routing logic legible in a way the composite headline score does not.
CAPABILITY CLUSTERS SHIFT WHERE DIFFERENTIATION CAN LIVE
When frontier models converge within measurement error on general capability, the headline number loses its power to justify a vendor choice. Differentiation — commercially durable differentiation — has to move somewhere else: price, latency, tooling integration, uptime reliability, or distribution reach. This is consistent with previously reported enterprise buying behaviour, not a prediction derived from this data. Nothing here claims these models are substitutable for any particular use case; that depends entirely on the task and on factors the ECI does not measure, including cost.









