The denominator test
Price per token, or per unit of completed work?
Share of spend, or share of tasks?
Investment added, or investment refinanced?
Four events, or one cause seen four times?
A probability, or a price on a small book?
A reading habit, not advice. It applies to this piece too.
Ten unrelated-looking AI stories from the week of September 16–22, 2026 turn out to be the same story told ten times: not about what AI can do, but about what our instruments can no longer see.
What Happened
Ten AI stories published on this site between September 16 and September 22, 2026 looked, on the surface, like unrelated dispatches: a model launch, a bond issue, a laptop pre-order, a European regulation, a prediction market move, a chip decision, a mathematics claim. Read back together, they describe a single structural problem. In each case the news was not really about what an AI system could do. It was about an instrument that could no longer measure what it was pointed at. Nobody need have acted in bad faith — every qualification that traveled with these stories at first publication travels with them here too.
The failures share a shape. Every one of them is, at bottom, a denominator problem: price per what, share of what, investment counted how, events counted how, a probability backed by what, a claim verified by whom. A price index assumes the good is roughly the same good between readings. An event count assumes events are independent. Peer review assumes claims arrive more slowly than they can be checked. Those assumptions were reasonable when they were made. They are what is now breaking — not the arithmetic.
The dangerous part is the failure mode. When the underlying thing moves faster than an instrument’s assumptions, the instrument does not stop working in any visible way. It keeps returning numbers that look entirely plausible. A gauge that reads zero gets replaced within the hour; a gauge that reads something reasonable gets quoted in a board paper. That is the actual risk across these ten stories, and it is a property of measurement under fast change rather than anything specific to artificial intelligence.
The key insight: The risk is not that these instruments were broken. It is that they kept returning plausible-looking numbers — numbers that would pass a board-paper smell test — while the denominator had silently moved. A broken gauge gets replaced. A reasonable-looking gauge gets cited.
The Ten Instruments — What Each Could Not See
Each instrument is named below alongside the specific thing it could not see. The qualifications that accompanied each story at first publication are preserved in full.
Posted Price (Grok 4.7) — could not see quality-adjusted deflation. Grok 4.7 held at $2 and $6 per million tokens, identical to 4.6, while the company’s own reported benchmarks rose. Those are vendor figures, not independent verification — and a flat sticker price on a more capable model is a real price cut that the sticker does not show.
Blended Token Price — could not separate a change in price from a change in mix. When relayed card-spend data showed frontier usage share and blended price both falling, the instrument could not tell you whether cheaper models were getting cheaper or whether the mix had shifted toward cheaper models. Those are different stories wearing the same number.
Running Totals of AI Capital (SoftBank) — could not separate new money from refinanced money. Per Reuters’ reading of a term sheet, SoftBank’s $10 billion of notes cancel an earlier $10 billion bridge funding the same commitment. Summing both figures produces a double-count rather than a new investment.
Sticker Price (Googlebook) — stopped being comparable once a term was bundled in. The Googlebook is on pre-order at $899 with twelve months of Google AI Pro included on every unit. The bundle has not been valued here. A sticker price that now contains a subscription is not the same unit as a sticker price that did not — the denominator changed.
Event Counts — mistook one root cause for four events. Several Irregular-linked disclosures arrived separately over weeks; every party disclosed. Google states the model believed the systems were in scope, that it stopped, and that it does not class the events as misalignment. No wrongdoing is alleged here. Counting sequential disclosures of one cause as independent events inflates the denominator of “how many times.”
The Front Month (Polymarket / Anthropic) — mistook a timing revision for a verdict revision. The Anthropic “by 31 October” contract fell from 54.5 to 5.2; the November contract rose from 22.75 to 67.65; the “by 31 December” contract moved only from 85.0 to 79.5. These are prices, not forecasts. Anthropic has not filed publicly, has not priced, and has not listed. A front-month price collapse in a thin contract is not a revised assessment of probability — it is a timing trade.
A Price Without a Book Size — mistook a thin market for a consensus. Two AI contracts showed liquidity of roughly $16.9 thousand and $8.2 thousand — thousands, not millions. A price derived from a book that small is not a market consensus. It is not an opportunity either. The denominator “how much conviction backs this price” was invisible in the headline number.
Unit-Count Comparison (DeepSeek / Chip Substitution) — understated a substitution penalty that scales worse than linearly. By Liang Wenfeng’s own estimate, the equivalent compute required approximately 50,000 Nvidia GB300 or approximately 200,000 Huawei Ascend 950 units — excluding pre-run experiments — after an earlier Ascend 910C run failed and the company reverted to Nvidia. A unit count without a performance-per-unit denominator understates the gap by a factor determined by that excluded denominator. No unlawful conduct is alleged.
A Rating Is Not a Limit (EU Data Centre Regulation) — conflated a proposed classification with a decided standard. The EU proposed a rating scheme for data centres above 500 kW. A separate consultation on minimum standards closes on 14 December 2026, with nothing decided. A proposed rating is not a cap, an obligation, or a ceiling. The denominator “how much of this is in force” was zero at time of publication.
Peer Review (OpenAI / Mathematics) — could not keep pace with a capability claim. OpenAI says an internal model resolved Navier-Stokes and more than 100 open problems. This is an unverified company claim about an internal model. Nothing here states that any problem is solved, that any prize was awarded, or that any advisory member endorsed the result. The claim has not entered the process that would resolve it — peer review — so peer review’s denominator is currently undefined.
The Structural Read
The Map of AI framework identifies nine layers of the AI stack — from raw compute through models, APIs, infrastructure, and applications — and the companies occupying each. One of the framework’s core uses is identifying which layer is generating the signal and which is generating the noise in any given week’s coverage. This week the signal is not at any single layer. It is at the measurement layer that sits above all of them: the layer of instruments, indices, counts, prices, and ratings that the industry uses to track itself.
The problem the ten instruments share is structural, not individual. Each was designed for a world in which the thing being measured changed slowly relative to the measurement cycle. A price index designed when model quality shifted quarterly does not break when quality shifts weekly — it just returns a number that looks correct but means something different. The index is still doing its arithmetic. The arithmetic’s relationship to the underlying reality has detached. The instrument is not malfunctioning. It is operating outside its design envelope.
The Map of AI — Measurement Layer
The instrument layer is the one nobody built on purpose
Every layer of the AI stack — compute, models, APIs, applications — has been consciously constructed and is actively contested. The measurement layer that tracks all of them was not designed for this environment. It was inherited from adjacent industries: finance, semiconductor manufacturing, academic publishing, consumer electronics. Each of those industries moves at a pace those instruments were built for. AI, in its current phase, does not. The ten instruments above are not outliers. They are the default state of the measurement layer right now.
There is also a selection caveat that the thesis itself demands. The ten stories below were chosen for coverage partly because they were interesting. A pattern assembled after the fact from stories one chose to write about is not evidence that the pattern is general. It is a hypothesis with ten illustrations, drawn from one publication’s own selection. The same denominator test applies to that claim: ten out of how many, selected by whom, and against what alternative.
91,000+ executives read Business Engineer for the AI strategy frameworks cited by ChatGPT, Claude, and Perplexity.
This is not investment advice. This synthesis draws on reporting published separately here; each figure’s full qualifications are in the linked piece, and no new facts or sources are introduced. The OpenAI mathematics results are that company’s own unverified claim about an internal model — nothing above states that the Navier–Stokes problem is solved, that any prize was awarded, or that any advisory group member reviewed or endorsed the claim. The Grok benchmark figures are SpaceXAI’s own reported numbers, not independent evaluations. Market figures are prices rather than forecasts, Anthropic has not filed publicly, priced or listed, and thin liquidity is not presented as an opportunity. On the test-environment disclosures, every party disclosed, and Google states the model believed the systems were in scope, that it stopped, and that it does not class the events as misalignment. The chip counts are Liang Wenfeng’s own estimate, excluding pre-run experiments. The European measure is a proposed rating scheme, not a cap, and Googlebook devices are on pre-order with no value assigned to the bundled subscription. Nothing above alleges unlawful conduct or bad faith by anyone, takes any position on any government, party, policy, export control or sanction, or predicts anything. The pattern described is drawn from this publication’s own selection of stories across one week, which is a selection rather than a sample — it is a hypothesis with ten illustrations, and the same denominator test applies to it.









