Nvidia, Meta, and xAI Are Winning the AI Efficiency Era — Here’s Why the Margin Is Moving Down the Stack

As reported by CNBC, with market analysis from Gavin Baker.

The AI race is shifting from biggest model to cheapest intelligence — and investor Gavin Baker argues that’s the most bullish possible outcome for infrastructure, not the death of AI value.

The Efficiency Shift — Key Signals

90%+

of all tokens forecast to come from open-weight models within 18–24 months (Peter Fenton / Benchmark)

90%+

estimated inference margins at frontier labs today — the number Baker says is most at risk if the shift materializes

18–24mo

Fenton’s window — possibly by year-end 2026 — for open-weight volume dominance

0

disagreement between bulls and bears that efficiency is happening — they only split on what it means

What Happened

A CNBC report published July 10 crystallized a structural shift the industry has been feeling for months: the AI competition is no longer about who can train the biggest model. The new battleground is routing, cost, and control — systems that decide which model to use, when, and what company data or tools to bring in. Smaller, task-tuned models are routinely outperforming large general-purpose ones on specific tasks, and hybrid local-plus-cloud deployment is moving from experiment to enterprise standard.

Peter Fenton of Benchmark put a number on the trajectory: he predicts more than 90% of all tokens will come from open-weight models within 18–24 months — possibly by year-end 2026. If that forecast lands anywhere near correct, frontier labs running 90%-plus inference margins would face real structural pressure, as companies run capable models without paying provider markups. The shift isn’t speculative noise; it’s the one data point AI bulls and bears have both co-signed.

That rare consensus is where the debate gets sharp. AI bear Michael Burry — who has publicly shorted the AI trade including Nvidia and derided what he calls “tokenmaxxing” as a bubble — read the efficiency shift as validation of his thesis: frontier infrastructure is overvalued if inference margins collapse. Investor Gavin Baker then quote-tweeted Burry to argue the precise opposite. Same shift, inverted conclusion. Baker’s read is the bull case worth unpacking carefully.

The key insight: Bulls and bears agree the AI efficiency shift is real. They disagree on causality. Baker’s argument is that cheaper intelligence doesn’t destroy AI value — it expands total token demand via the Jevons dynamic, and redistributes margin dollars from the model layer to the infrastructure layer. Burry reads the same facts as bubble deflation. The structural question is which causal chain is right — and by Baker’s own admission, the answer isn’t in yet.

How the Efficiency Debate Unfolded

Early 2026

Task-tuned small models begin outperforming frontier models on specific enterprise benchmarks; open-weight adoption accelerates across cloud and on-prem deployments

Mid-2026

xAI releases Grok 4.5 at Opus-class capability but significantly lower pricing, putting direct pressure on frontier-model pricing power; Meta continues aggressive open-source model releases

July 10, 2026

CNBC publishes the efficiency-shift thesis; Fenton’s 90%-open-weight-tokens forecast goes wide; Michael Burry cites it as bubble confirmation

July 10–13, 2026

Gavin Baker quote-tweets Burry: same efficiency shift, opposite conclusion — mega bull case for AI infrastructure, not a bust; the debate bifurcates across the investment community

The Structural Read

Baker’s argument has a precise internal logic, and it’s worth laying out exactly — because it’s easy to misread as simple AI optimism. It isn’t. It’s a specific claim about where in the stack value accrues when the model layer commoditizes.

The mechanism runs like this: if frontier-lab inference margins compress from 90%+ toward something lower — because open-weight or cheap closed models take volume share — the per-unit cost of intelligence falls. Cheaper intelligence means better ROI for end customers. Better ROI means companies deploy AI into more workflows, more tokens, more volume. The Jevons paradox applied to compute: efficiency doesn’t shrink the market, it expands total demand. Margin dollars don’t disappear; they redistribute. They flow away from the model provider and toward whoever has the lowest per-token infrastructure cost at scale. That is, structurally, Nvidia’s toll-booth position.

Baker also names the vertically integrated players — SpaceX/xAI and Meta — as structurally well-positioned: they own both the model and the distribution, run cheaper models that are competitive on capability, and don’t need to extract margin at the model layer because they monetize elsewhere. Jensen Huang’s consistent push for open source, Baker argues, is rational self-interest: lower model-layer margin percentages mean more inference compute overall, and Nvidia captures a toll on all of it.

Gavin Baker — Critical Hedge

“Cheap, mostly open-source tokens are likely the majority of volume today, but the majority of economic value still accrues to the most intelligent models. Might change. We will see.”

That hedge is load-bearing. Baker is presenting a conditional thesis, not a realized fact. The Jevons dynamic and the margin-migration story are the would-be consequence if open-weight volume dominance materializes and cheap tokens drive enterprise deployment at scale. Right now, by his own read, the economic value in AI still pools at the frontier. The thesis is a forecast about where the gravity shifts — not a description of where it has already shifted.

AI Value Chain

When Capability Commoditizes, the Toll Booth Wins

The AI Value Chain framework maps where margin accrues at each layer of the stack. When a layer commoditizes — as model capability appears to be doing — value doesn’t evaporate. It migrates to the layer that controls the scarce resource. In the current architecture, that’s compute infrastructure at scale. The router layer is the next candidate. See the full framework: The AI Value Chain →

Three Implications

IMPLICATION #1 — THE MARGIN MOVES TO INFRA

If open-weight volume dominance materializes, frontier labs face compressed inference margin percentages — but total token demand expands via the Jevons dynamic, so margin dollars don’t vanish; they flow to whoever runs the cheapest compute at scale. That is Nvidia’s structural position, and it’s why Jensen’s open-source advocacy is rational strategy, not altruism. The Nvidia toll-booth dynamic gets stronger, not weaker, if cheap inference explodes volume. Treat this as a conditional: it’s the outcome if the shift Fenton forecasts arrives at scale.

IMPLICATION #2 — THE ROUTER IS THE NEW BATTLEGROUND

The CNBC framing is precise: the winning AI product is “a system that decides which model to use, when, and what tools or company data to bring in.” That is routing and orchestration — and it is already the central strategic bet at Perplexity (model-agnostic orchestration) and the logic behind Nadella’s “Choice” principle at Microsoft. The router doesn’t need to win the model race; it needs to compose cheap and smart models more efficiently than anyone else. When capability is abundant and cheap, the orchestration layer captures the coordination premium. This is the Routing Paradigm for Enterprise AI playing out in real time.

IMPLICATION #3 — VERTICAL INTEGRATION IS POSITIONED, BUT THE TIMING IS EVERYTHING

Baker singles out xAI and Meta as structurally well-placed: vertically integrated, strong cheaper models, alternative monetization that doesn’t depend on inference margin. Grok 4.5’s Opus-class capability at lower pricing is a live example of the commoditize-the-frontier move. But Baker’s own hedge is the critical constraint: today, economic value still pools at the most intelligent models. Vertically integrated players positioned for the efficiency era are early — which is either a competitive moat being built in advance, or a bet that may not pay off on the timeline assumed. The bull case is real; it just hasn’t been realized yet.

How the AI Stack Reprices (If the Shift Lands)

Scroll to Top

Discover more from FourWeekMBA

Subscribe now to keep reading and get access to the full archive.

Continue reading

FourWeekMBA