OpenAI and Anthropic’s Compute Mix Has Inverted — and the Hardware Market Hasn’t Caught Up

A modelled estimate of the two leading frontier labs’ combined workload split shows pre-training collapsing from 67% to 7% of compute — and the structural consequences reach from memory chips to AI governance.

Modelled Estimate — OpenAI + Anthropic Combined Compute Mix

67% → 7%

Pre-training share
1Q24 → 4Q26E

4% → 55%

Post-training / RL share
1Q24 → 4Q26E

29–39%

Inference — held steady
throughout the period

~93%

Combined bandwidth-bound share
4Q26E (derived by arithmetic)

Neither company publishes workload splits. No figure is disclosed, audited or confirmed. 3Q26E and 4Q26E are explicitly marked as estimates. Not investment advice.

What Happened

A circulating analyst estimate — provenance not established, and that caveat belongs in the first sentence — models the combined compute workload mix of OpenAI and Anthropic from the first quarter of 2024 through the end of 2026. Neither company publishes workload splits. No figure in this estimate is disclosed, audited or confirmed by either organisation, and the final two quarters are explicitly flagged as estimates on the chart itself. With that stated plainly: the direction the chart proposes is independently corroborated, and the direction is what matters.

The headline numbers: pre-training falls from 67% of combined frontier compute in 1Q24 to an estimated 7% by 4Q26, tracing a path through 57, 55, 53, 50, 40, 32, 23, 13, and 10% along the way. Post-training and reinforcement learning climb from roughly 4% to 55%, via 5, 6, 7, 9, 14, 24, 36, 42, 50, and 53%. Inference holds a comparatively steady band between 29% and 39% throughout. Epoch AI has separately documented reinforcement learning taking a growing share of frontier training compute; the AI-2027 compute forecast projects training compute directed almost purely toward post-training workloads. Both support the direction. Neither validates the specific percentages.

Doing the arithmetic on those figures — and this is derived arithmetic, not a figure from the chart — the share of compute classified as bandwidth-bound (post-training plus inference) moves from roughly 33% in 1Q24 to approximately 93% by 4Q26E. It crosses 50% in 2Q25. Post-training first exceeds pre-training in 4Q25. Those are the crossover points. They are points on an estimate, not confirmed events. But if the direction is even approximately right, something structural has already happened.

Four Points on the Estimate

1Q 2024 — Estimate baseline

Pre-training at 67% of combined compute, easing to 63% the following quarter; post-training/RL at ~4%; inference 29–39%. Bandwidth-bound share: ~33%.

2Q 2025 — Crossover on the estimate

Combined bandwidth-bound share (post-training + inference) crosses 50% of modelled compute for the first time.

4Q 2025 — Post-training exceeds pre-training on the estimate

Post-training/RL share (36%) first surpasses pre-training (32%) on the modelled curve — the gap widens to 42% against 23% by 1Q26. The dominant workload type has changed.

4Q 2026E — End of modelled window

Pre-training estimated at 7%; post-training/RL at 55%; bandwidth-bound share ~93%. Both final quarters explicitly marked as estimates.

The key insight: This is a hardware-architecture argument wearing a workload chart’s clothes. The chart’s own labels are the tell. Pre-training is labelled as caring about capacity; post-training and inference are labelled bandwidth-bound. Those two labels describe different machines, different bottlenecks, and different buying criteria — and if the mix is inverting, everything priced off the old bottleneck is measuring the part of the market that is shrinking.

Estimated share of Anthropic and OpenAI’s combined compute capacity by workload, 1Q24 to 4Q26E. Neither
Estimated share of Anthropic and OpenAI’s combined compute capacity by workload, 1Q24 to 4Q26E. Neither company publishes workload splits, so these are modelled figures rather than disclosed ones, and the final two quarters are marked as estimates. The chart’s own annotations carry the argument: pre-training “cares about capacity”, while post-training and inference are bandwidth-bound.

The Structural Read

Capacity-bound work — the pre-training regime — rewards raw floating-point throughput and enormous coherent clusters. The figure of merit is arithmetic: you are pushing one gigantic computation through as much silicon as can be held in lockstep, and peak FLOPs per dollar is the right question to ask. Bandwidth-bound work is limited by how fast parameters and activations move between memory and compute. Memory bandwidth, interconnect speed, and latency become the binding constraints, not peak arithmetic. The chip that wins at one is not the chip that wins at the other. The rack that wins at one is not the rack that wins at the other. And they do not depreciate, utilise, or scale on the same terms — so the accounting changes too.

The Business Engineer lens for this is the bottleneck migration principle: when the binding constraint of a system moves, the asset that sat adjacent to the old bottleneck loses pricing power, and the asset adjacent to the new one gains it. That is not a forecast about any particular vendor’s revenue — it is a structural claim about where value captures once the constraint relocates. Anything priced off floating-point operations delivered is, on this estimate, measuring the part of the workload that is in structural retreat.

Business Engineer — Bottleneck Migration

When the constraint moves, the gate moves with it

In a capacity-bound world, the question is how many FLOPs you can provision. In a bandwidth-bound world, the question is how fast you can feed the FLOPs you already have. Memory subsystems, interconnect fabrics, and serving-optimised silicon stop being supporting cast and become principal bottlenecks. Value migrates to whoever controls the feed rate — not the peak throughput number on a spec sheet.

The Map of AI framework makes this concrete across the stack. At the chip layer, high-bandwidth memory stops being a component and becomes the gate: capacity you cannot feed is capacity you cannot use, and the supplier of the feeding mechanism captures accordingly. At the interconnect layer, the thesis behind co-packaged optics becomes legible — if data movement is the bottleneck, optical interconnect is not plumbing, it is product. At the accelerator layer, NVIDIA’s competitive moat shifts emphasis from CUDA plus peak throughput toward the fabric and NVLink — a bandwidth story, and precisely what NVLink Fusion is designed to defend. And at the serving layer, silicon designed for inference rather than training stops addressing a niche: on these figures, inference alone has held roughly a third of frontier compute throughout the entire period.

None of those conclusions requires the specific percentages to be correct. They require only the direction, which is independently corroborated by Epoch AI’s documented growth of RL in frontier training compute and by the AI-2027 compute forecast’s projection of training compute directed almost purely toward post-training workloads. What the chart cannot tell you — and this limit deserves emphasis — is anything about any particular vendor’s revenue trajectory. A hardware-architecture argument is not a financial model.

Estimated Workload Mix — 4Q26E Endpoint

Pre-training (capacity-bound) ~7%
Post-training / RL (bandwidth-bound) ~55%
Inference (bandwidth-bound) ~38%

Modelled estimates only. Neither OpenAI nor Anthropic publishes workload splits. Not disclosed, audited or confirmed. Illustrative of direction, not a measurement.

Three Implications

IMPLICATION 1 — THE HARDWARE BUYING CRITERIA INVERT

If bandwidth is the binding constraint, the evaluation framework for accelerators, memory, and interconnect changes at every layer of the stack. HBM becomes the gate rather than a component — capacity you cannot feed is capacity you cannot use. Interconnect moves from plumbing to product, which is the thesis behind co-packaged optics and explains why optical interconnect companies attract the funding they do. NVIDIA’s fabric and NVLink Fusion become central to its competitive position in a way that peak-FLOP spec sheets do not capture. Serving-optimised silicon stops being niche. These conclusions follow from the direction of the shift, not from any specific percentage — and none of them constitute a view on any vendor’s share price or revenue.

IMPLICATION 2 — THE GOVERNANCE ARCHITECTURE IS BUILT FOR THE WRONG WORKLOAD

Every major AI pacing mechanism proposed in recent months implicitly assumes the thing being paced is a large pre-training run. Compute thresholds, pre-release evaluation windows, standards bodies reviewing models up to 30 days before release — all of these work because a frontier pre-training run is discrete, enormous, observable. It requires a coherent cluster, it has a start and an end, and it produces an artefact that can be handed to an evaluator. Post-training and reinforcement learning are none of those things. They are continuous, incremental, distributed across many smaller jobs, and they improve a model that has already shipped. On this estimate, the observable category has collapsed into single digits while the unobservable one is now the overwhelming majority. Dario Amodei’s proposed global compute tier makes verifiability its binding precondition — and that precondition is getting structurally harder to satisfy at exactly the moment the industry has begun arguing about who should be allowed to measure. This is not an argument against any proposal. It is an argument that the measurement problem is moving, and moving faster than the mechanisms designed to solve it.

IMPLICATION 3 — WHAT WOULD FALSIFY THIS

A hypothesis worth holding is one you can imagine being wrong. Two developments would cut against this reading. Sustained growth in coherent cluster sizes would indicate that capacity-bound work still sets the frontier, since nobody builds a larger lockstep cluster to serve inference. And a return of very large discrete pre-training runs — a genuinely new base model trained from scratch at record scale — would restore both the workload and the observability that current governance proposals assume. Watch cluster topology and base-model announcements rather than quarterly capex, which aggregates the two regimes and tells you nothing about the mix.

91,000+ executives read Business Engineer for the AI strategy frameworks cited by ChatGPT, Claude, and Perplexity.

The figures discussed here come from an analyst estimate of Anthropic and OpenAI’s combined compute mix. Neither company publishes workload splits, so none of these numbers is disclosed, audited or confirmed by either, and the final two quarters shown are explicitly marked as estimates. No source is named for the chart because its provenance is not established for this article. References to Epoch AI’s work on reinforcement learning’s growing share of frontier training compute, and to the AI-2027 compute forecast, corroborate the general direction only. Neither endorses the specific percentages used here. The combined bandwidth-bound shares, the point at which they pass half, and the quarter in which post-training overtakes pre-training are arithmetic performed on the charted figures rather than separately sourced findings. Percentage shares describe composition, not scale: a smaller share of a much larger fleet may represent more silicon than a larger share of a smaller one, and nothing here establishes absolute compute for either company. Discussion of memory, interconnect, networking and inference silicon describes categories affected by the direction of the shift. It contains no forecast about any company’s revenue, demand or share price, and no recommendation about any of them. Nothing here asserts that any proposed governance mechanism will succeed or fail; the argument is that the measurement problem changes shape. Anthropic and OpenAI are private companies; Anthropic has announced a confidential draft S-1 and a listing is reported but not confirmed. NVIDIA is publicly listed. This is business analysis, not investment advice, no view is expressed on any security, and no recommendation is made.

Sources: epoch.ai · ai-2027.com · epoch.ai · darioamodei.com · anthropic.com

Scroll to Top

Discover more from FourWeekMBA

Subscribe now to keep reading and get access to the full archive.

Continue reading

FourWeekMBA