Artificial Analysis and ARC Prize supply the empirical other shoe to OpenAI’s “AGI era” launch — and the scorecard is more modest, and more expensive, than the framing.
What Happened
A day after OpenAI launched GPT-6 Astra and president Greg Brockman closed the briefing with “Welcome to the AGI era,” the first independent post-launch measurements arrived — and they describe a more sober model than the launch did. Artificial Analysis published its Intelligence Index (version 4.1.1, a nine-evaluation composite, dated September 3) showing GPT-6 Astra at its highest effort settings scoring approximately 61. That places it dead level with the model it replaces — GPT-5.6 Sol, also at 61 — and behind Claude Fable 5.1 at 66 and Claude Opus 5 at 63, while costing roughly two and a half times what Sol costs: $10 per million input tokens and $50 per million output versus Sol’s $4 and $20.
Separately, the ARC Prize Foundation — a genuinely independent second measurement body — tested OpenAI’s most dramatic launch claim: 98.6% on the ARC-AGI-3 reasoning benchmark. ARC’s own standard-harness result came in near 62.7%. The near-perfect figure reproduces only under a different harness configuration — a provider-adapter setup — rather than ARC’s standard evaluation environment. ARC itself characterizes the ARC-AGI-3 result as a real generational jump over prior models; the dispute is about which number gets headlined, not whether progress occurred.
Two guardrails matter before drawing conclusions. First, Artificial Analysis is, so far, the sole source of the intelligence-index ranking. Trending Topics, OfficeChai, and Implicator are all downstream of the same Artificial Analysis data — they are not independent corroborations, and the AA numbers are approximately 24–48 hours old. This is one benchmark house plus ARC Prize, not broad independent consensus. Second, “Astra trails its rivals” would overclaim: Muse Spark 1.3’s 62 on the AA index is a partner-only preview “max” variant; on publicly available settings it ties Astra at approximately 61. Mythos 5.1 carries no independent AA score at all because access remains restricted, so any Anthropic edge there is a vendor claim, not an independently verified result. And there is no verified Astra entry on LMArena or SWE-bench to cite. None of this is investment advice.
AA Intelligence Index — Where Astra Actually Lands
Source: Artificial Analysis Intelligence Index v4.1.1, Sep 3, 2026. AA is the sole source of this index; downstream outlets draw on the same data. Muse Spark 1.3 publicly available ~61; Mythos 5.1 unscored (restricted access).
The key insight: On the one independent general-intelligence index available, GPT-6 Astra is a flat generation at a premium price — matching its own predecessor’s score while costing 2.5 times as much. Every prior model generation offered more capability for the same or less money. Astra inverts that deal. That inversion is the story, not the benchmark rank itself.
The Structural Read
This is the empirical other shoe to the launch, landing exactly where the launch-day analysis said to look. The launch-day piece flagged the gap between the “AGI era” rhetoric and the careful, gated rollout mechanics, and argued the mechanics were the honest signal. The independent numbers now supply the capability version of that same gap — on two separate dimensions.
The price is the story. Flat capability at 2.5 times the cost inverts the generational deal the industry has run on since GPT-3. Every prior generation offered a clear proposition: more capability for the same or less money. That proposition is what justified enterprise re-platforming, developer adoption, and the underlying assumption that switching costs would rise with each cycle. Astra breaks that pattern on general intelligence. A flat-capability, higher-price release is a pricing-power move dressed as a capability leap — and sophisticated enterprise buyers will price against the independent scorecard, not the launch deck.
The harness-dependent headline. OpenAI led with 98.6% on ARC-AGI-3. ARC Prize’s standard-harness result: approximately 62.7%. The gap — more than thirty points — is not evidence of bad faith. It is a demonstration that a flagship benchmark number can be harness-dependent, and that the number a lab chooses to headline and the number an independent body reproduces under its own standard conditions can differ materially. This is the durable lesson of the week, and it generalizes: OpenAI’s other launch numbers — DeepSWE, OSWorld, FrontierMath, ExploitBench — remain unreplicated vendor claims that independent bodies have not yet tested. Treat them accordingly.
The vendor-versus-independent gap is now a pricing signal. The correct frame is narrow and defensible: on general intelligence, Astra is flat at a premium; its flagship benchmark is harness-dependent; and it is competitive-to-improved on agentic coding (AA Coding Agent Index: ~67, roughly level with Opus 5, with Fable 5.1 near 70 — a real gain at lower cost than Fable). ARC itself calls the ARC-AGI-3 result a genuine generational jump. So “Astra is bad” is wrong. “Astra is a flat-capability general-intelligence release at a premium price with a harness-dependent headline benchmark and a genuine coding gain” is what the independent evidence supports. That is a materially different proposition than “the AGI era.”
BE Framework — Vendor-vs-Independent Gap
The Inverted Generational Deal
When a model generation delivers flat general-intelligence scores at a higher price point, the launch narrative carries more weight than the capability delta. That structural shift — rhetoric running ahead of independent measurement — is visible across the week: the same framing that drew a congressional bill to ban superintelligence on launch day is the framing independent benchmarks just declined to confirm. The marketing is running ahead of both the measurement and the politics, and both are starting to catch up. See the full week’s through-lines: Five Through-Lines. And the congressional response: Ban Superintelligence Act.
This is business analysis, not investment advice. The Intelligence Index ranking comes from a single benchmark house (Artificial Analysis); ARC Prize is a separate independent body. Several figures cited by OpenAI at launch remain unreplicated vendor claims. Benchmark results are recent and may be revised.









