OpenAI’s AGI Benchmark Has a Denominator — and That Changes How to Read the 80% Claim

Alex Heath’s TIME reporting surfaced a concrete internal benchmark alongside the percentage — and the benchmark is the more important object.

What Happened

Writing for TIME after more than two weeks and interviews with over twenty OpenAI executives, employees, investors and rivals, Alex Heath reported that Chief Research Officer Mark Chen estimates the company is “80% of the way” to AGI. Sam Altman’s formulation was more careful and worth carrying precisely: by the end of the year, he said, OpenAI would have an internal system he would call AGI. That qualifier — he would call — is part of the actual claim, not a softener someone added later.

The detail that separates this from a percentage floating in air is what Chief Scientist Jakub Pachocki told TIME: a system called Astra already hits an internal benchmark. The benchmark is statable in a single sentence — hand it an experimental idea, and it writes the code in OpenAI’s codebase, runs the experiment, and reports back with results. That is the denominator against which “80%” is coherent. Co-founder Greg Brockman believes people will look back on this moment as the one when AGI emerged; that is his prediction, attributed to him, and it is not adopted here.

What the remaining twenty per cent consists of, how the eighty per cent was arrived at, and whether it represents a formal internal measure or an estimate offered in conversation are not established in the reporting and do not appear below. Neither does any published OpenAI definition of AGI, any benchmark score, or any capability of Astra beyond the described task.

The key insight: A percentage requires a denominator, and here two different denominators are wearing the same three-letter word. Against Pachocki’s stated benchmark the figure is a coherent engineering statement about a target a third party could in principle check. Against what most readers hear in “AGI” there is no denominator at all — because there is no agreed operational definition to divide by. Both can be true simultaneously, and holding both is the entire analytical discipline.

The number is precise. The useful question is what it is a percentage of.
The number is precise. The useful question is what it is a percentage of.

The Structural Read

Credit is due on a point that usually goes unremarked: OpenAI published the denominator. Pachocki described the benchmark in concrete, falsifiable terms rather than gesturing at capability in the abstract. An internal target that can be stated in one sentence and examined by an outsider is a far better object than a vague aspiration, and most organisations claiming progress toward something ambitious never offer one at all.

The problem is not the source. The problem is transmission. A figure that is rigorous at origin can arrive three repetitions later as something close to meaningless, without anyone having been inaccurate at any point. The definition is reliably the first thing to fall off in the relay — for an unglamorous reason: it is the longest part of the claim and the least quotable. The number is four characters and travels perfectly. The sentence explaining what it measures does not.

This publication has encountered the same structural shape three times in the past week: a run rate that annualised a single month, a contracted total quoted with no term attached, and a figure that was simultaneously a run rate, pro forma, and conditional on a deal that had not closed. None of those involved dishonesty either, and neither does this one. No motive is attributed to anyone here — the gap is not dishonesty, it is one word carrying two definitions. They involved one number doing two jobs, and an audience with no way to tell which job had arrived. That is the pattern worth naming.

Product Overhang Doctrine — Business Engineer

Capability Builds Before the Label Settles

The Product Overhang Doctrine holds that capability accumulates invisibly, then surfaces all at once — usually before the vocabulary for it is ready. What Pachocki’s benchmark describes is an engineering milestone with a specific task definition. What the word “AGI” carries is decades of science-fiction, philosophy, and contested academic usage. When a measurable milestone and an unmeasured cultural concept share the same label, the gap between them is not a gap in honesty — it is a vocabulary problem that precision alone cannot close. The benchmark is the more durable object. It will still be examinable when the debate over the label has moved on.

Sam Altman — via Alex Heath, TIME

“By the end of the year it would have an internal system he would call AGI.”

Altman’s own qualifier deserves to survive intact. “A system he would call AGI” is explicitly a statement about his own classification — not a claim that the field, a regulator, or a philosopher would agree with him. That qualifier is sitting in his sentence, doing real work. It survives only if everyone repeating the claim carries it forward.

Three Implications

THE BENCHMARK IS THE DURABLE OBJECT

Pachocki’s task description — take an experimental idea, write the code, run it, report back — is falsifiable and examinable by an outsider. That is a higher standard than most ambitious capability claims reach. As the debate over what “AGI” means continues to evolve, the benchmark will remain checkable long after the label debate has shifted. Analysts and observers tracking OpenAI’s progress have something concrete to update against; that is unusual and worth acknowledging.

TRANSMISSION DEGRADES PRECISION, NOT INTENT

The four-link chain — source to reporter to podcast to reader — reliably strips definition and preserves number. No one in that chain needs to misrepresent anything for the original precision to be lost. This is a structural feature of how technical claims travel through media, not a pathology of this story in particular. The implication for anyone building on this reporting: go back to Heath’s TIME piece, not a summary of a summary, and carry Altman’s qualifier and Pachocki’s benchmark forward together.

HOLD THE PRECISION LIGHTLY, NOT DISMISSIVELY

Whether “80%” represents a formal internal measure or a considered numerical estimate offered in conversation is not established. A percentage implies a measured scale; a judgement expressed numerically is a different kind of object — and offering one in an interview with TIME is an entirely normal executive behaviour. That distinction is a reason to hold the figure’s precision lightly rather than a reason to doubt Chen. Both the figure and the benchmark deserve to be taken seriously on their own terms.

Business Engineer Framework

The Map of AI — Where Benchmarks Sit in the Stack

Pachocki’s benchmark describes a capability at the research-infrastructure layer of the AI stack — the layer where models meet internal tooling, codebases, and experimental pipelines. The Map of AI maps 200+ companies across nine layers and shows why progress at this layer has different downstream consequences than progress at the model or application layer. Understanding which layer a claim lives on changes what it means for everyone else in the stack.

Explore the Map of AI →

The Bottom Line

The most analytically useful thing in Alex Heath’s TIME reporting is not the percentage — it is the sentence Jakub Pachocki used to describe the benchmark. That sentence has a subject, a verb, and a falsifiable outcome. It is the denominator “80%” needs to mean anything, it is the object that survives when the label debate moves on, and it is the thing most people repeating the claim will not carry with them. Read the number through the benchmark, carry Altman’s qualifier intact, and the reporting holds. Drop either one and you are left with a number doing two jobs — which is how a rigorous figure becomes, over four links, something close to noise.

Sources: Alex Heath, TIME — Sam Altman / OpenAI interview; The Decoder — Altman on AGI by end of 2026. Structural analysis: Business Engineer / FourWeekMBA editorial. Published September 26, 2026. Not investment advice.

91,000+ executives read Business Engineer for the AI strategy frameworks cited by ChatGPT, Claude, and Perplexity.

A denominator is stated. The eighty per cent is progress toward an internal benchmark that Jakub Pachocki described concretely — a system that takes an experimental idea, writes the code in OpenAI’s codebase, runs the experiment and reports back. That is narrower than what the phrase AGI conveys in general use, and both facts are carried above. Sam Altman’s own formulation was a system he would call AGI, which is a statement about what he would call it rather than a claim the field would agree. Nothing above says OpenAI is close to AGI, or that it is not. No motive is attributed to Mark Chen, Sam Altman, Jakub Pachocki or Greg Brockman, and nothing above suggests anyone is redefining the term to claim success. Greg Brockman’s view that this period will be remembered as the moment AGI emerged is his belief and a prediction, and is not adopted here. What the remaining twenty per cent consists of, how the eighty per cent was arrived at and whether it is a formal measure are not established; a considered estimate offered in an interview is a normal thing for an executive to give and nothing above suggests the figure was invented. Astra’s capabilities beyond the described benchmark, any benchmark score, evaluation or release date, any published OpenAI definition of AGI, and other executives’ or rivals’ views are not established and do not appear above. Nothing above is investment advice or predicts anything.

Scroll to Top

Discover more from FourWeekMBA

Subscribe now to keep reading and get access to the full archive.

Continue reading

FourWeekMBA