Skip to content
KO EN
AI 인프라

NVIDIA’s CUDA Monopoly Is Cracking, but the Real Moat Lies Deeper in the Fabric Layer

NVIDIA's (NVDA) dominance in AI chips is facing a structural challenge that goes beyond mere market share fluctuations. Cerebras CEO Andrew Feldman recentl

Editorial illustration for AI infrastructure & semiconductor coverage

NVIDIA’s (NVDA) dominance in AI chips is facing a structural challenge that goes beyond mere market share fluctuations. Cerebras CEO Andrew Feldman recently estimated on the Mad Podcast that CUDA has lost roughly 70% of frontier AI training share over the past 24 months. This is not a speculative claim; it is supported by public evidence: Google’s Gemini trains exclusively on its own TPUs, Anthropic’s Claude runs on AWS Trainium, and Google has begun offering TPUs to external customers for the first time. Even OpenAI, still primarily on CUDA, is co-developing custom silicon. The era of a single dominant training substrate is ending.

What Happened: The Rise of Multi-Silicon Training

Feldman’s estimate is grounded in verifiable shifts. Google confirmed in 2024 that Gemini trains on in-house TPUs, a program dating back to 2016. Amazon’s Trainium chips, designed to reduce AWS’s reliance on NVIDIA, are now the primary training infrastructure for Anthropic’s Claude, as announced through AWS partnerships. These are not experimental deployments; they represent the primary compute for two of the most prominent frontier model families. However, as Gennaro Cuofano argues in Beyond NVIDIA’s Moat, framing this as ‘CUDA losing share’ misses the structural point. The real story lies one layer deeper: the fabric interconnect.

Why It Matters: The Fabric Layer Is the True Moat

Frontier model training is not a single-chip problem. Models like GPT-4 and Gemini 1.5 require splitting computations across thousands of accelerators that must constantly synchronize state. The bottleneck is not raw compute but the speed of data movement between chips. This is why NVIDIA acquired Mellanox for $6.9 billion in 2019, gaining end-to-end control of the distributed training stack through NVLink (chip-to-chip) and InfiniBand (node-to-node). When developers praised CUDA’s ecosystem, they were actually describing the compounding advantage of the fabric layer underneath. Google and Amazon have now built their own: Google’s TPU pods use ICI (Inter-Core Interconnect), and Amazon’s Trainium clusters use EFA (Elastic Fabric Adapter), directly challenging NVIDIA’s fabric monopoly.

Our Analysis: CUDA’s Decline, Fabric’s Resilience

XPLAIN AI interprets this not as NVIDIA’s downfall but as market diversification. CUDA’s loss of frontier training share is real, but it reveals that NVIDIA’s core strength was always the fabric, not the software. NVIDIA still holds the industry’s best fabric technology in NVLink and InfiniBand, which remain critical for inference and enterprise AI workloads. Google and Amazon’s custom chips are optimized within their own cloud ecosystems, and it remains unclear whether they can match NVIDIA’s performance for external customers. The competitive axis may shift from ‘where you train’ to ‘where you infer,’ a market where NVIDIA’s fabric advantage still holds strong.

Potential Winners and Risks

  • Potential beneficiaries: Google (GOOGL) and Amazon (AMZN) strengthen their AI semiconductor ecosystems, reducing NVIDIA dependency. AMD (AMD) and Intel (INTC) may find new opportunities in a less CUDA-dominated market. AI chip startups like Cerebras and Graphcore could gain traction in frontier training.
  • Potential risks: NVIDIA (NVDA) faces erosion of its monopoly in frontier training, though near-term impact is limited by growing inference and enterprise demand. TSMC (TSM) benefits from customer diversification but could see short-term volume shifts if NVIDIA orders decline.

Counter-Scenario and Uncertainties

Feldman’s 70% estimate should be weighed against his company’s competitive position against NVIDIA. Google and Amazon’s custom chips may still lag in generality and software ecosystem maturity. OpenAI’s custom silicon is in development and unlikely to fully replace CUDA soon. NVIDIA could also widen the gap with next-generation fabric technologies like NVLink 6 or Quantum 2 InfiniBand. The fabric moat is not yet breached; it is merely being tested.

Key Metrics to Watch

Investors should monitor: (1) Google and Amazon’s AI semiconductor CAPEX trends; (2) the share of NVIDIA’s data center revenue from frontier AI customers like OpenAI and Meta; (3) real-world adoption of competing chips like AMD’s MI300X or Intel’s Gaudi 3; and (4) NVIDIA’s new fabric-related contracts and technology announcements. These indicators will reveal whether NVIDIA’s fabric moat holds or competitors are closing the gap.

#AISemiconductors #NVIDIA #GoogleTPU #AWSTrainium #CUDA #FabricLayer #AITraining #ChipCompetition

Sources

Written by: XPLAIN AI Editorial Team · Reviewed by: XPLAIN AI Editorial Desk
This content was drafted with AI assistance based on publicly available sources and reviewed under XPLAIN AI's editorial standards.

Found an error? Request a correction →