The Eight-Layer AI Infrastructure Stack
The AI ecosystem forms an hourglass: layers narrow toward the HBM chokepoint at Layer 5, then widen again below. Click any layer to expand its details.
Demand flows down
Supply flows up
The Bottleneck Shift: From Compute to Memory
The AI infrastructure constraint has evolved through distinct phases, each defined by a different binding limitation on scaling.
Phase 1: GPU Shortage (2020-2023)
Phase 2: Memory Shortage (2024-2028+)
2020
AI scaling era begins; GPU shortage starts
2021
H100 announced; demand surges
2022-23
GPUs scarce at any price; TSMC at capacity
2024
Constraint shifts to HBM; memory shortage emerges
2025
HBM demand at 58% CAGR; H200 ships 192GB
2026
HBM4 expected; Micron Hiroshima breaks ground
2027
Capacity expansions ramp; shortage persists
2028+
Micron Hiroshima ships; structural shortage continues
The HBM Oligopoly: Three Companies Control AI Scaling
Only three companies in the world can manufacture HBM at scale. Barriers to entry (DRAM expertise, TSV stacking, advanced packaging, capital intensity) make new entrants effectively impossible.
SK Hynix
Market leader; closest NVIDIA relationship; first-mover in HBM3e
Samsung
Strong volume; yield challenges in latest-gen HBM
Micron
Smallest but fastest-growing; targeting 20%+ share
Geopolitical concentration risk: 88% of HBM production is in South Korea. Micron's $9.6B Hiroshima investment (with ~$3.2B Japanese gov't subsidy) positions Japan as the second global HBM pillar. China is locked out by export controls.
Key Metrics: The Numbers Behind the Chokepoint
0 TB/s
Peak HBM3e Bandwidth
vs. traditional DRAM: 10-100x improvement
0 GB
H200 HBM Capacity
Up from 80GB on H100
0% CAGR
HBM Demand Growth
Through 2030 -- steepest in semiconductors
0-50 $/GB
HBM Cost per GB
vs. $3-5 for standard DRAM (10x premium)
0%
HBM Share of GPU Cost
Memory is the asset; compute is the wrapper
0 B$
Micron Hiroshima Investment
~$3.2B in Japanese gov't subsidies
The Demand Flywheel: Every AI Vendor Needs HBM
HBM demand is driven by the full spectrum of AI accelerator vendors, creating a multi-source demand flywheel that outpaces GPU demand growth.
NVIDIA
Dominant consumer. H100 (80GB), H200 (192GB), B100 (192GB+). NVIDIA's roadmap defines HBM demand curves.
AMD
MI300X packs 192GB HBM. Credible alternative for memory-intensive workloads. Growth translates directly to HBM demand.
Google TPU
Custom TPU accelerators are HBM-intensive, serving internal AI workloads and Google Cloud customers.
Custom Silicon
AWS Trainium, Meta MTIA, Microsoft Maia all incorporate HBM. Hyperscaler custom chips add incremental demand.