A single set of model weights carries four verified ARC Prize scores at once — and that fact exposes something structural about what benchmark numbers actually measure.
What follows is analysis of published benchmark results. It is not investment advice, and it predicts nothing about AGI, benchmark saturation, model releases or which laboratory wins.
What Happened
On 24 July 2026, ARC Prize published verified results for Anthropic’s Claude Opus 5 across three benchmark generations. The same weights, on the same date, produced four distinct scores: 97.5% on ARC-AGI-1 at Max reasoning effort, 90.4% on ARC-AGI-2’s semi-private set at Max, 88.3% on ARC-AGI-2’s public evaluation set at High, and 30.16% on ARC-AGI-3’s public demonstration set at High. None of those numbers contradicts any other. Each is a precise, fully sourced result — and each belongs to a different triple of model, harness, and evaluation set.
Separately — and this is drawn from secondary coverage rather than any primary Nvidia document — Nvidia’s AVO agent is reported to have scored 100% on ARC-AGI-3’s public demonstration set, the same set that produced Claude Opus 5’s 30.16%. AVO is reported to run mainly on Claude Opus 5 as its underlying model. Françis Chollet, as reported, called the approach “very nice” while noting the run covered the tutorial-like public set and did not extend to ARC-AGI-3’s semi-private or private sets. No primary Nvidia document on AVO’s architecture, cost, latency, or compute has been established, and no semi-private or private ARC-AGI-3 score for AVO exists in the published record.
A Y Combinator Paper Club episode on self-improving harnesses brought both numbers into the same sentence. The host — unnamed here because the episode description does not name them — described the 30% figure as the best result on “the holdout on the private that no one else has access to other than Greg and Chollet,” then contrasted it with AVO’s 100%. The 30.16% is real and matches the ARC Prize record exactly. The set description is a speaking error: ARC Prize records it as ARC-AGI-3 Public Demo, not a closely held private holdout. The host was speaking extempore; no motive is imputed and no dishonesty is alleged.
The key insight: Because both the 30.16% and the 100% sit on the same public demonstration set, the host’s comparison is actually valid — and by mislabelling the set as a private holdout, he made his own point sound weaker than it is. Checking provenance rescued the argument, not debunked it.

The Structural Read
A benchmark score is not a property of a model. It is a property of a triple: the model, the harness wrapped around it, and the evaluation set it was scored on. Drop any one element and the number does not become imprecise — it becomes unusable. The four verified Claude Opus 5 figures make this concrete rather than pedantic. The same weights are simultaneously a 97.5% system and a 30.16% system, and there is no fact of the matter about which is “the real score” because there is no such thing as the real score. There is only a score on a set under a harness.
None of this is a criticism of ARC Prize. The discrepancy examined here is checkable precisely because ARC Prize publishes all three parts of the triple correctly and dates them. The problem surfaces when any one element is dropped in transmission — which is what happened in the podcast, and what routinely happens whenever a single percentage circulates without its full context.
Harness Theory · Business Engineer
“Once scaffolding moves a score further than the weights do, the product boundary and the measurement boundary stop agreeing. A leaderboard organised by model name is quietly reporting on the wrong object.”
The AVO case illustrates this precisely. The reported gap between 30.16% and 100% on the same public demonstration set is not a contest between two laboratories — it sets the same underlying weights against themselves. The entire reported distance is accounted for by the scaffolding: a loop of inspection, planning, implementation, and evaluation with persistent memory and supervision, as described in secondary coverage. The host’s own characterisation was “this thing that doesn’t deserve any research, just some wrapper and some scaffolding.” On this set, the wrapper accounts for nearly the whole distance.
There is also a second property worth keeping from the provenance error. The intuition about mislabelled evaluation sets is that they flatter: someone quotes a score from an easy public split and lets the audience assume it came from the hard private one. Here the opposite happened. Describing a public-demo score as a private holdout made a same-set comparison look like a cross-set comparison — precisely the kind of claim that should be discarded on sight. Checking which set produced the number rescued the comparison rather than disqualifying it. A provenance error is not a directional bias. It is a loss of information, and lost information can cut either way.
Three Implications
IMPLICATION 1 · The Triple Is the Unit of Measurement
Communicating a benchmark score as a single number — without naming the harness and the evaluation set — is not a simplification. It removes the information that makes the number meaningful. The four verified Claude Opus 5 figures are the clearest demonstration of this in the published ARC Prize record: the same weights, on the same date, span from 30.16% to 97.5% depending on which set and effort level you specify. Any claim that compares two scores is implicitly claiming that the harness and the set are held constant. That claim needs to be made explicit, not assumed.
IMPLICATION 2 · Scaffolding Is a Product Layer, Not an Asterisk
If the reported gap between base-model performance and harness-augmented performance on the same public set is close to 70 percentage points, the scaffolding is not a footnote to the model result — it is most of the result. That has a direct consequence for how AI products are categorised and competed on. A leaderboard ranked by model name, when the dominant variable is the harness, is measuring the wrong object. Competitive advantage built on scaffolding is real and reproducible, but it accrues to the harness builder, not necessarily to the model provider.
IMPLICATION 3 · Provenance Errors Are Symmetric, Not Flattering by Default
The standard concern about benchmark provenance is inflation: an easy set presented as a hard one. This case ran the other direction — a public set described as a private holdout, which made a valid same-set comparison look methodologically suspect and appeared to undermine the speaker’s own argument. The discipline of checking which set produced a number is not a sceptic’s reflex for knocking results down. It is the only way to know what direction the error runs. On this occasion, it confirmed the comparison rather than disqualifying it.
91,000+ executives read Business Engineer for the AI strategy frameworks cited by ChatGPT, Claude, and Perplexity.
The four Claude Opus 5 figures above — 97.5% on ARC-AGI-1 at Max, 90.4% on ARC-AGI-2 Semi-Private at Max, 88.3% on ARC-AGI-2 Public Eval at High, and 30.16% on ARC-AGI-3 Public Demo at High — are ARC Prize verified results dated 24 July 2026. The 100% for Nvidia’s AVO, the statement that AVO runs mainly on Claude Opus 5, and François Chollet’s caveat are as reported in secondary coverage and are not drawn from a primary Nvidia document. The mislabelling of the evaluation set is treated here as an extempore speaking error on a podcast. No motive is imputed, no dishonesty is alleged, and the remark is not characterised as misleading or as a misrepresentation. The host is not named because the episode description does not name them. Seth Karten’s 95.5% figure for Prime Agent is an unverified speaker claim whose evaluation set is not established here, and it is deliberately not placed alongside the verified results. Any Nvidia primary document on AVO, AVO’s architecture, cost, latency or compute, any semi-private or private ARC-AGI-3 score for AVO or Prime Agent, Anthropic comment, other laboratories’ ARC scores and ARC-AGI-3’s task count or human baseline are not established and do not appear. Nothing above predicts AGI, benchmark saturation, model releases or which laboratory wins, and nothing above is investment advice.
Sources: arcprize.org · youtube.com · thenewstack.io · ARC Prize verified results for Claude Opus 5, 24 July 2026 (supplied primary) · YC Paper Club clip, supplied verbatim; secondary coverage of Nvidia AVO (supplied as reported)









