A modelled estimate of the two leading frontier labs’ combined workload split shows pre-training collapsing from 67% to 7% of compute — and the structural consequences reach from memory chips to AI governance.
What Happened
A circulating analyst estimate — provenance not established, and that caveat belongs in the first sentence — models the combined compute workload mix of OpenAI and Anthropic from the first quarter of 2024 through the end of 2026. Neither company publishes workload splits. No figure in this estimate is disclosed, audited or confirmed by either organisation, and the final two quarters are explicitly flagged as estimates on the chart itself. With that stated plainly: the direction the chart proposes is independently corroborated, and the direction is what matters.
The headline numbers: pre-training falls from 67% of combined frontier compute in 1Q24 to an estimated 7% by 4Q26, tracing a path through 57, 55, 53, 50, 40, 32, 23, 13, and 10% along the way. Post-training and reinforcement learning climb from roughly 4% to 55%, via 5, 6, 7, 9, 14, 24, 36, 42, 50, and 53%. Inference holds a comparatively steady band between 29% and 39% throughout. Epoch AI has separately documented reinforcement learning taking a growing share of frontier training compute; the AI-2027 compute forecast projects training compute directed almost purely toward post-training workloads. Both support the direction. Neither validates the specific percentages.
Doing the arithmetic on those figures — and this is derived arithmetic, not a figure from the chart — the share of compute classified as bandwidth-bound (post-training plus inference) moves from roughly 33% in 1Q24 to approximately 93% by 4Q26E. It crosses 50% in 2Q25. Post-training first exceeds pre-training in 4Q25. Those are the crossover points. They are points on an estimate, not confirmed events. But if the direction is even approximately right, something structural has already happened.
The key insight: This is a hardware-architecture argument wearing a workload chart’s clothes. The chart’s own labels are the tell. Pre-training is labelled as caring about capacity; post-training and inference are labelled bandwidth-bound. Those two labels describe different machines, different bottlenecks, and different buying criteria — and if the mix is inverting, everything priced off the old bottleneck is measuring the part of the market that is shrinking.

The Structural Read
Capacity-bound work — the pre-training regime — rewards raw floating-point throughput and enormous coherent clusters. The figure of merit is arithmetic: you are pushing one gigantic computation through as much silicon as can be held in lockstep, and peak FLOPs per dollar is the right question to ask. Bandwidth-bound work is limited by how fast parameters and activations move between memory and compute. Memory bandwidth, interconnect speed, and latency become the binding constraints, not peak arithmetic. The chip that wins at one is not the chip that wins at the other. The rack that wins at one is not the rack that wins at the other. And they do not depreciate, utilise, or scale on the same terms — so the accounting changes too.
The Business Engineer lens for this is the bottleneck migration principle: when the binding constraint of a system moves, the asset that sat adjacent to the old bottleneck loses pricing power, and the asset adjacent to the new one gains it. That is not a forecast about any particular vendor’s revenue — it is a structural claim about where value captures once the constraint relocates. Anything priced off floating-point operations delivered is, on this estimate, measuring the part of the workload that is in structural retreat.
Business Engineer — Bottleneck Migration
When the constraint moves, the gate moves with it
In a capacity-bound world, the question is how many FLOPs you can provision. In a bandwidth-bound world, the question is how fast you can feed the FLOPs you already have. Memory subsystems, interconnect fabrics, and serving-optimised silicon stop being supporting cast and become principal bottlenecks. Value migrates to whoever controls the feed rate — not the peak throughput number on a spec sheet.
The Map of AI framework makes this concrete across the stack. At the chip layer, high-bandwidth memory stops being a component and becomes the gate: capacity you cannot feed is capacity you cannot use, and the supplier of the feeding mechanism captures accordingly. At the interconnect layer, the thesis behind co-packaged optics becomes legible — if data movement is the bottleneck, optical interconnect is not plumbing, it is product. At the accelerator layer, NVIDIA’s competitive moat shifts emphasis from CUDA plus peak throughput toward the fabric and NVLink — a bandwidth story, and precisely what NVLink Fusion is designed to defend. And at the serving layer, silicon designed for inference rather than training stops addressing a niche: on these figures, inference alone has held roughly a third of frontier compute throughout the entire period.
None of those conclusions requires the specific percentages to be correct. They require only the direction, which is independently corroborated by Epoch AI’s documented growth of RL in frontier training compute and by the AI-2027 compute forecast’s projection of training compute directed almost purely toward post-training workloads. What the chart cannot tell you — and this limit deserves emphasis — is anything about any particular vendor’s revenue trajectory. A hardware-architecture argument is not a financial model.
Estimated Workload Mix — 4Q26E Endpoint
Modelled estimates only. Neither OpenAI nor Anthropic publishes workload splits. Not disclosed, audited or confirmed. Illustrative of direction, not a measurement.
Three Implications
IMPLICATION 1 — THE HARDWARE BUYING CRITERIA INVERT
If bandwidth is the binding constraint, the evaluation framework for accelerators, memory, and interconnect changes at every layer of the stack. HBM becomes the gate rather than a component — capacity you cannot feed is capacity you cannot use. Interconnect moves from plumbing to product, which is the thesis behind co-packaged optics and explains why optical interconnect companies attract the funding they do. NVIDIA’s fabric and NVLink Fusion become central to its competitive position in a way that peak-FLOP spec sheets do not capture. Serving-optimised silicon stops being niche. These conclusions follow from the direction of the shift, not from any specific percentage — and none of them constitute a view on any vendor’s share price or revenue.
IMPLICATION 2 — THE GOVERNANCE ARCHITECTURE IS BUILT FOR THE WRONG WORKLOAD
Every major AI pacing mechanism proposed in recent months implicitly assumes the thing being paced is a large pre-training run. Compute thresholds, pre-release evaluation windows, standards bodies reviewing models up to 30 days before release — all of these work because a frontier pre-training run is discrete, enormous, observable. It requires a coherent cluster, it has a start and an end, and it produces an artefact that can be handed to an evaluator. Post-training and reinforcement learning are none of those things. They are continuous, incremental, distributed across many smaller jobs, and they improve a model that has already shipped. On this estimate, the observable category has collapsed into single digits while the unobservable one is now the overwhelming majority. Dario Amodei’s proposed global compute tier makes verifiability its binding precondition — and that precondition is getting structurally harder to satisfy at exactly the moment the industry has begun arguing about who should be allowed to measure. This is not an argument against any proposal. It is an argument that the measurement problem is moving, and moving faster than the mechanisms designed to solve it.
IMPLICATION 3 — WHAT WOULD FALSIFY THIS
A hypothesis worth holding is one you can imagine being wrong. Two developments would cut against this reading. Sustained growth in coherent cluster sizes would indicate that capacity-bound work still sets the frontier, since nobody builds a larger lockstep cluster to serve inference. And a return of very large discrete pre-training runs — a genuinely new base model trained from scratch at record scale — would restore both the workload and the observability that current governance proposals assume. Watch cluster topology and base-model announcements rather than quarterly capex, which aggregates the two regimes and tells you nothing about the mix.
91,000+ executives read Business Engineer for the AI strategy frameworks cited by ChatGPT, Claude, and Perplexity.
The figures discussed here come from an analyst estimate of Anthropic and OpenAI’s combined compute mix. Neither company publishes workload splits, so none of these numbers is disclosed, audited or confirmed by either, and the final two quarters shown are explicitly marked as estimates. No source is named for the chart because its provenance is not established for this article. References to Epoch AI’s work on reinforcement learning’s growing share of frontier training compute, and to the AI-2027 compute forecast, corroborate the general direction only. Neither endorses the specific percentages used here. The combined bandwidth-bound shares, the point at which they pass half, and the quarter in which post-training overtakes pre-training are arithmetic performed on the charted figures rather than separately sourced findings. Percentage shares describe composition, not scale: a smaller share of a much larger fleet may represent more silicon than a larger share of a smaller one, and nothing here establishes absolute compute for either company. Discussion of memory, interconnect, networking and inference silicon describes categories affected by the direction of the shift. It contains no forecast about any company’s revenue, demand or share price, and no recommendation about any of them. Nothing here asserts that any proposed governance mechanism will succeed or fail; the argument is that the measurement problem changes shape. Anthropic and OpenAI are private companies; Anthropic has announced a confidential draft S-1 and a listing is reported but not confirmed. NVIDIA is publicly listed. This is business analysis, not investment advice, no view is expressed on any security, and no recommendation is made.
Sources: epoch.ai · ai-2027.com · epoch.ai · darioamodei.com · anthropic.com









