Analysis draws on Cerebras CEO Andrew Feldman’s remarks on the Mad Podcast and Gennaro Cuofano’s “Beyond NVIDIA’s Moat” (The Business Engineer). Where frontier labs train (Gemini/Claude/OpenAI) corroborated by public reporting linked below.
Cerebras CEO Andrew Feldman estimates CUDA has lost ~70% of frontier training share in 24 months. The structural reason that number isn’t a death sentence for NVIDIA lives one layer deeper.
What Happened
On a recent episode of the Mad Podcast, Cerebras CEO Andrew Feldman made an observation worth sitting with: two years ago, essentially every state-of-the-art AI model was trained on NVIDIA GPUs inside a CUDA workflow. Today, Google’s Gemini trains on Google’s own TPUs, Anthropic’s Claude trains on AWS Trainium, and Google has begun making TPUs available to external customers — a first. Of the major US frontier labs, OpenAI still trains primarily on CUDA, though it is actively co-developing custom silicon. Feldman’s estimate is that CUDA has lost roughly 70% of frontier training share across that 24-month window. He is careful to present this as his estimate; the directional shift is not seriously in dispute.
The public record supports the direction. Google’s TPU program is well-documented going back to its 2016 first-generation announcement; the decision to use TPUs exclusively for Gemini training was confirmed through Google DeepMind communications and infrastructure disclosures. Amazon’s Trainium chips — designed specifically to reduce AWS’s dependence on NVIDIA for AI training workloads — have been publicly tied to Anthropic’s training infrastructure through AWS’s own partnership announcements. These are not experimental deployments; they are the primary substrate for two of the three most-cited frontier model families.
In Beyond NVIDIA’s Moat, Gennaro Cuofano of The Business Engineer argues that treating this as “CUDA losing share” misframes the structural question. The more useful analysis starts one layer down — at the fabric.
The key insight: CUDA was never really the moat. It was the visible surface of a moat whose foundation is distributed-compute networking — NVLink, InfiniBand, and the systems engineering that ties thousands of accelerators into a single coherent training run. That layer, not the programming language, is what NVIDIA’s 2019 Mellanox acquisition was actually purchasing.
The Structural Read
Frontier model training is not a single-chip problem. A training calculation for a model at GPT-4 or Gemini 1.5 scale is far too large to fit on one accelerator, so it is split across thousands of chips that must constantly share state to remain consistent — what engineers call tensor model parallelism. The bottleneck is not raw compute; it is how fast information moves between those chips. If the interconnect fabric is slow or unreliable, additional compute is literally worthless.
This is why NVIDIA paid $6.9 billion for Mellanox in 2019. NVLink (chip-to-chip within a node) and InfiniBand (node-to-node across a data center) gave NVIDIA end-to-end ownership of the distributed training stack. CUDA held because the fabric held underneath it. When developers said “NVIDIA’s software ecosystem is a moat,” they were largely describing the compounding advantage of having the only interconnect fabric optimized for this specific workload at scale.
Google and Amazon have now built vertically-integrated alternatives: TPU pods with their own high-bandwidth interconnects (Google’s ICI fabric), and Trainium clusters with EFA (Elastic Fabric Adapter) networking. These are not NVIDIA-compatible — and that is the point. They are closed, optimized stacks that allow their owners to train frontier models without touching NVIDIA’s supply chain or licensing any NVIDIA software. The CUDA layer has been circumvented at the top two hyperscalers. What has not been circumvented is NVIDIA’s continuing dominance of the open-market, multi-tenant, and enterprise training infrastructure where no single operator has the scale to justify building a custom fabric.
Gennaro Cuofano — Beyond NVIDIA’s Moat, The Business Engineer
“The CUDA moat held as long as it did because the fabric moat sat underneath it. NVIDIA’s durable training advantage was never really CUDA-the-programming-language. Frontier training is a distributed-systems problem — if you cannot move information between accelerators fast enough, raw compute is worthless.”
The two-sided read, then: NVIDIA’s infrastructure position is secure for the foreseeable future. It remains the volume leader by a wide margin, with the only broadly-available distributed-compute ecosystem (CUDA plus NVLink plus InfiniBand) that third-party enterprises can actually buy and operate. What is loosening is long-term control at the frontier training layer specifically. The training tier has become structurally multi-silicon: TPU and Trainium as credible frontier-scale alternatives, a second tier including Cerebras wafer-scale chips, AMD’s MI-series, and Broadcom-designed custom ASICs for the next-largest labs, and a Chinese branch built around Huawei Ascend operating entirely outside US export controls.
There is one more layer to this: NVIDIA’s support for the Open-Weight Alliance reads differently once you understand the fabric logic. NVIDIA does not want AI value and pricing power to concentrate at the model layer — a world dominated by two or three proprietary closed models would allow those labs to extract margin from NVIDIA on hardware procurement. Vigorous competition at the model layer, including open-weight models that commoditize the model itself, keeps the industry structurally dependent on the infrastructure layer NVIDIA still dominates. The Open-Weight Alliance is partly a product of genuine open-source conviction; it is also entirely consistent with NVIDIA’s infrastructure lock-in strategy.
NVIDIA Moat — Layer-by-Layer Status
Fabric Layer (NVLink / InfiniBand)
INTACTNo third-party alternative at comparable scale and availability. The Mellanox acquisition compounds over time; distributed training at enterprise scale still runs through NVIDIA’s fabric.
CUDA Ecosystem (software gravity)
MIXEDDominant for inference, mid-tier labs, and enterprise. Losing ground exclusively at the frontier training layer where hyperscalers have built closed vertical alternatives.
Frontier Training Monopoly
ERODINGFeldman estimates ~70% share loss at frontier in 24 months. Google (TPU) and Amazon/Anthropic (Trainium) have exited. OpenAI remains on CUDA but is co-developing custom silicon.
Inference + Enterprise Volume
DOMINANTThis is where NVIDIA’s revenue is actually concentrated. Inference deployments, cloud GPU instances, and enterprise AI are not migrating to custom silicon at meaningful scale.
Three Implications
FOR HYPERSCALERS: VERTICAL INTEGRATION IS NOW THE FRONTIER NORM
Google and Amazon have demonstrated that frontier-scale training on custom silicon is not a moonshot — it is operational. Any hyperscaler still routing frontier training through NVIDIA is now making a deliberate strategic choice, not an engineering default. The cost and supply-chain leverage of owning the full stack — chip, fabric, compiler, orchestration — at this scale justifies the years of investment it requires.
FOR ENTERPRISES: THE INDEPENDENCE INTEGRATOR POSITION STRENGTHENS
If the training substrate is now structurally multi-vendor — TPU, Trainium, CUDA, MI-series, Ascend — no single infrastructure provider can hold an enterprise captive through compute lock-in alone. The defensible enterprise position, as the Independence Integrator thesis argues, is to own the routing and governance junction and treat compute as a rentable commodity. Abstraction layers that run identically across substrates are now a genuine product category, not a theoretical one.
FOR INVESTORS: THE MOAT QUESTION IS A LAYER QUESTION
Evaluating NVIDIA’s competitive position by asking “is CUDA losing share?” is the wrong unit of analysis. The correct question is: which layer of the distributed training stack is contestable, and over what time horizon? The fabric layer — NVLink, InfiniBand, the systems-level Mellanox acquisition — remains effectively uncontested outside hyperscaler-owned closed stacks. NVIDIA’s revenue risk is concentrated in a potential future where inference also migrates to custom silicon at scale, which is a 5-to-10-year problem, not a 2-year one.









