NVIDIA vs. Custom ASICs: GPU Dominance Persists Despite Startup Performance Claims · history
Version 11
2026-07-05 02:08 UTC · 142 items
What
NVIDIA reportedly holds approximately 92% of the GPU market [2] while Jensen Huang publicly characterized competitor custom ASICs as 'science projects' compared to NVIDIA's AI factory platform [10]. Etched exited stealth with $800M raised, over $1B in customer contracts, and 400+ engineers recruited from NVIDIA, TSMC, and other semiconductor companies [13][14], claiming 10x inference improvement over incumbents via a cluster-scale memory architecture [15][16] — with no independent benchmarks published. OpenAI's Jalapeño inference chip, built in nine months with Broadcom, targets gigawatt-scale deployment with Microsoft starting in 2026 [19][21], also without external validation. The central question is whether CUDA's software depth — illustrated by a 2.5x serving cost reduction in under 70 days [4] — is a structural advantage custom silicon challengers cannot close.
Why it matters
Etched's launch with substantial funding, customer commitments, and deep industry talent, alongside OpenAI's Jalapeño deployment, provide two structurally distinct tests of whether custom inference silicon can challenge NVIDIA's software moat in production: a hardware-first startup and a hyperscaler with captive workloads. Whether either converts benchmark claims into verified production results will determine whether the software moat thesis holds.
Open questions
Etched has $800M in funding and over $1B in customer contracts [13], but claims 10x inference improvement over incumbents [16] without published independent benchmarks — does the cluster-scale memory architecture perform in production at that scale?
Can Jalapeño's claimed better performance-per-watt and lower serving costs survive independent benchmarking, given no external validation has been published? [19][21]
Does the 2.5x serving cost reduction on NVIDIA's GB200 NVL72 via CUDA kernel rewrites in under 70 days [4] represent a rate of software improvement that inference-specific ASICs cannot match on near-term timelines?
Do hyperscalers with captive workloads — OpenAI, Meta, Amazon — constitute a meaningfully different structural class than hardware-first startups in their ability to develop the software stack required to make custom silicon viable? [3][23][24][19]
Narrative
NVIDIA has held or grown AI compute market share against custom ASICs through mid-2026 [1], reportedly commanding approximately 92% of the GPU market [2]. The analytical explanation with the most traction centers on software: SemiAnalysis argues that AI chip competition is decided by software ecosystem depth, not hardware benchmarks, and that 99% of custom ASIC projects fail in production because building a competitive software stack is the actual hard problem — one NVIDIA has largely solved over two decades [3]. A concrete data point: serving costs on NVIDIA's GB200 NVL72 dropped 2.5x in under 70 days through CUDA kernel rewrites applied to the Kimi model architecture [4], illustrating software-layer improvement that hardware-first competitors cannot replicate on short timelines. NVIDIA has simultaneously structured custom silicon as a participant in its own ecosystem via NVLink Fusion, spanning Marvell and SiFive (chip designers), Samsung Foundry (manufacturing), Astera Labs (connectivity), and MediaTek [5][6][7][8][9].
Jensen Huang made his sharpest public statement on competition in early July 2026, characterizing competitor custom ASICs as 'science projects' compared to NVIDIA's AI factory platform, adding that competitors are still copying NVIDIA's previous generation while its roadmap is at the physical limits of semiconductor scaling, and that NVIDIA is 'the only platform to build on' when betting hundreds of billions on AI infrastructure [10]. This framing directly responds to a wave of challenger claims: Tensordyne announced 13x rack throughput versus NVIDIA's NVL72 GB300 on DeepSeek-R1 [11], and DeepAdapt claimed 82% AI operating cost reduction by shifting GPU workloads to CPUs [12] — both based on internal simulations without independent validation, fitting the pattern SemiAnalysis describes.
Etched exited stealth on June 30, 2026 with $800M raised, over $1B in customer contracts, and a team of 400+ engineers recruited from NVIDIA, TSMC, and other leading semiconductor companies [13][14]. The company claims 10x inference improvement over incumbent chipmakers via a cluster-scale memory (CSM) architecture that creates a shared low-latency memory pool across an entire scale-up domain, combining HBM and SRAM to target both throughput and interactive inference speeds [15][16]. Andrej Karpathy publicly praised the approach's extreme low-voltage engineering targeting tokens-per-watt at interactive speeds [17]; SemiAnalysis noted the design's high SerDes count [18] without assessing production viability. No independent benchmarks have been published. OpenAI and Broadcom announced Jalapeño on June 24 — OpenAI's first custom LLM inference chip, built in nine months — claiming better performance per watt and lower serving costs, targeting gigawatt-scale deployment with Microsoft starting in 2026 [19][20][21]. SemiAnalysis noted the nine-month timeline requires 'making no mistakes' [22].
The structural argument distinguishing hyperscalers from startups runs as follows: Meta expanding its MTIA silicon program [23] and Amazon entering the AI chip market [24] bring captive workloads and engineering depth that most hardware-first startups lack — the same workload ownership SemiAnalysis identifies as the key variable determining whether custom silicon can develop a competitive software stack. OpenAI brings that same structural advantage with Jalapeño. Etched's $1B+ in customer contracts and deep industry talent [13][14] represent the strongest version of the startup counterargument, but production performance rather than funding or headcount will determine whether the software moat thesis holds.
Timeline
- 2026-01-15: SiFive announces integration of NVLink Fusion for next-generation RISC-V AI data center chips. [8]
- 2026-03: Meta announces expansion of its MTIA custom silicon program to power next-generation AI workloads. [23]
- 2026-06-14: NVIDIA CEO Jensen Huang states publicly that next year's growth will remain rapid. [25]
- 2026-06-17: Tensordyne announces an inference rack claiming 13x throughput versus NVIDIA's NVL72 GB300 on DeepSeek-R1, based on internal simulations. [11][31]
- 2026-06-18: Milk Road AI publishes analysis arguing NVIDIA has held or grown market share against ASICs, contrary to two years of consensus predictions. [1]
- 2026-06-19: SemiAnalysis states 99% of custom ASIC projects fail and that AI chip success is determined by software depth, not hardware specifications. [3]
- 2026-06-19: Jensen Huang appears at a Marvell keynote framing NVLink Fusion as enabling custom silicon coexistence within NVIDIA-led clusters. [26]
- 2026-06-19: Amazon reported to be entering the AI chip market; DeepAdapt claims 82% AI operating cost reduction by shifting GPU workloads to CPUs, without independent validation. [24][12]
- 2026-06-21: Samsung Foundry added to NVIDIA's NVLink Fusion ecosystem as a manufacturing partner for custom silicon. [6]
- 2026-06-22: SemiAnalysis documents 2.5x serving cost reduction on GB200 NVL72 in under 70 days via CUDA kernel rewrites. [4]
- 2026-06-24: OpenAI and Broadcom announce Jalapeño, OpenAI's first custom LLM inference chip built in nine months, claiming better performance per watt; deployment with Microsoft targets gigawatt scale starting in 2026. [19][20][32][27]
- 2026-06-24: SemiAnalysis notes the Jalapeño nine-month development timeline requires 'making no mistakes.' [22]
- 2026-06-30: Etched exits stealth with $800M raised, over $1B in customer contracts, and 400+ engineers from NVIDIA and TSMC; claims 10x inference improvement via cluster-scale memory architecture; Karpathy praises the low-voltage design. [17][18][15][16][13][14]
- 2026-07-03: Jensen Huang calls competitor custom ASICs 'science projects' relative to NVIDIA's AI factory platform, characterizing NVIDIA as the only viable platform for large-scale AI infrastructure bets. [10]
Perspectives
NVIDIA / Jensen Huang
Characterizes competitor custom ASICs as 'science projects' while NVIDIA builds 'revenue-generating AI factories'; says NVIDIA delivers a complete, industry-standardized platform and is the only viable option for large AI infrastructure bets, with competitors copying its previous generation while its roadmap is at semiconductor physical limits.
Evolution: Moved from bullish growth projections to direct public dismissal of custom ASIC competition; NVLink Fusion simultaneously positions custom silicon makers as ecosystem participants rather than rivals.
SemiAnalysis
AI chip competition is decided by software ecosystem depth; 99% of ASIC projects fail in production; Jalapeño's nine-month timeline requires 'making no mistakes'; Etched's CSM design has a high SerDes count.
Evolution: Consistent skeptic of hardware-first ASIC narratives; the June 2026 CUDA cost-reduction data sharpens the software moat argument; technical commentary now extends to both Etched's architecture and OpenAI's development timeline.
OpenAI / Broadcom
Jalapeño is OpenAI's first custom LLM inference chip, built in nine months with Broadcom; claims better performance per watt and lower serving costs than current accelerators, targeting gigawatt-scale deployment with Microsoft starting in 2026.
Evolution: Announced June 24; no independent benchmarks published; secondary coverage frames it as NVIDIA dependency reduction.
Etched
Claims 10x inference improvement over incumbent chipmakers via cluster-scale memory architecture; raised $800M with over $1B in customer contracts and 400+ engineers from NVIDIA, TSMC, and other semiconductor leaders.
Evolution: Exited stealth June 30 with explicit performance claims and substantial commercial traction; drew endorsement from Karpathy and architectural commentary from SemiAnalysis; no independent benchmarks published.
Milk Road AI
Data shows NVIDIA has held or grown market share against ASICs; NVLink Fusion makes custom silicon makers ecosystem participants rather than competitors.
Evolution: Consistently contrarian on ASIC displacement; reframes competition as coexistence via NVLink Fusion.
NVLink Fusion ecosystem (Marvell, SiFive, Samsung Foundry, MediaTek, Astera Labs)
Chip designers, foundries, and connectivity silicon makers operate within NVIDIA's interconnect framework rather than against it.
Evolution: Expanded from Marvell alone to a set spanning manufacturing, RISC-V AI chips, consumer chip design, and connectivity silicon.
Meta / Amazon (hyperscalers)
Both treating custom silicon as critical for scaling AI workloads — Meta expanding MTIA, Amazon entering the AI chip market — with captive workloads and engineering scale that differs structurally from hardware-first startups.
Evolution: Deepening commitment through 2026; structurally distinct from startups in their ability to develop accompanying software stacks alongside custom hardware.
Tensordyne / DeepAdapt
Both claim large performance advantages over NVIDIA-centric inference (13x throughput and 82% cost reduction respectively) based on internal simulations without independent validation.
Evolution: Fit the benchmark-slide pattern SemiAnalysis describes; neither claim has been independently verified.
Tensions
- Jensen Huang calls custom ASICs 'science projects' and says NVIDIA is the only viable AI infrastructure platform; Etched claims 10x inference improvement at its formal launch, backed by $800M raised, $1B+ in customer contracts, and 400+ engineers from NVIDIA and TSMC. [10][15][16][13][14][17]
- SemiAnalysis argues the CUDA software moat is real and deepening — 2.5x serving cost reduction in 70 days via kernel rewrites — while OpenAI claims Jalapeño delivers better performance per watt through workload-specific hardware design; SemiAnalysis's 'making no mistakes' comment signals skepticism of OpenAI's timeline. [4][19][22]
- SemiAnalysis argues AI chip success is determined by software depth and 99% of ASIC projects fail in production; Etched, OpenAI, Tensordyne, and DeepAdapt all publish performance or architectural claims framed as improvements over NVIDIA hardware, none with independent validation. [3][16][19][11][12]
- Milk Road AI and NVIDIA argue the 'NVIDIA versus everyone' frame is wrong — NVLink Fusion makes custom silicon a coexistent participant; SemiAnalysis supports NVIDIA's resilience through a different mechanism: software moat, not ecosystem embrace. [26][1][3]
- The software moat argument holds that startups cannot overcome CUDA's depth; Meta, Amazon, and OpenAI — organizations with captive workloads and engineering scale — test whether sufficient resources can develop the required software stack alongside custom hardware. [3][23][24][19]
Sources
- [1] Everyone assumed Nvidia would get crushed by ASICs but the data says the opposite just happened and the reason why chang… — Milk Road AI Twitter (2026-06-18)
- [2] Nvidia dominates the GPU market with a 92% market share; how... - moomoo Community — reactive:asic-gpu-market-dynamics
- [3] 100% of AI chip startups have slides/“simulated performance data” showing that their chip is way better, but 99% of cust… — SemiAnalysis Twitter (2026-06-19)
- [4] CUDA MOAT ALERT 🔥: In less than 70 days, GB200 NVL72 serving costs decreased by 2.5x through software improvements alone… — SemiAnalysis Twitter (2026-06-22)
- [5] In 2016, Marvell's largest design win was a Wi-Fi chip in the Barbie Dream House (Save this). — Milk Road AI Twitter (2026-06-20)
- [6] NVIDIA Adds Samsung Foundry to NVLink Fusion Ecosystem for Custom Silicon Manufacturing : r/hardware — reactive:asic-gpu-market-dynamics
- [7] Why Connectivity is the New Frontier of AI Infrastructure—and What NVLink Fusion Means for the Future — reactive:asic-gpu-market-dynamics
- [8] SiFive To Integrate Nvidia NVLink Fusion For Next-Gen AI Data Centers — reactive:asic-gpu-market-dynamics
- [9] MediaTek | NVLink Fusion | Custom AI ASIC Innovation — reactive:asic-gpu-market-dynamics
- [10] Jensen Huang on 'Nvidia vs. custom ASIC' — Rohan Paul Twitter (2026-07-03)
- [11] Quite a massive inferencing rack breakthrough from @TensordyneInc . — Rohan Paul Twitter (2026-06-17)
- [12] DeepAdapt has launched a runtime intelligence layer that cuts AI operating costs by up to 82% and 33X faster inference b… — Rohan Paul Twitter (2026-06-19)
- [13] Etched Emerges From Stealth With Working Chip, $800M — reactive:asic-gpu-market-dynamics
- [14] Etched Pulls 400+ Engineers From NVIDIA, TSMC & More to Build a New Frontier Inference Cluster For AI Which Is Already Worth $1B in Demand — reactive:asic-gpu-market-dynamics
- [15] Etched has come out of stealth with its first AI inference chip. And this isn't a normal launch. — reactive:asic-gpu-market-dynamics (2026-07-02)
- [16] Etched @Etched came out of stealth claiming it can beat incumbent chipmakers on inference not by 10% but by 10x. — reactive:asic-gpu-market-dynamics (2026-06-30)
- [17] Congrats!! I was impressed to learn about some of the engineering wizardry (e.g. *very* low voltage domains, cluster sca… — Andrej Karpathy Twitter (2026-06-30)
- [18] etched cluster-scale memory has so many SerDes https://t.co/V6yTAtsYf9 https://t.co/OZHcQshvUj — SemiAnalysis Twitter (2026-06-30)
- [19] OpenAI and Broadcom unveil LLM-optimized inference chip — OpenAI Blog (2026-06-24)
- [20] OpenAI and Broadcom announce chip designed for LLM inference at scale — Ars Technica AI (2026-06-24)
- [21] 🟡 Hollywood's AI trick work — Semafor Technology (2026-06-26)
- [22] Chat develop a chip from initial design to tape out in 9 months, make no mistakes. https://t.co/DaNCCPJebw — SemiAnalysis Twitter (2026-06-24)
- [23] Expanding Meta's Custom Silicon to Power Our AI Workloads — reactive:asic-gpu-market-dynamics
- [24] 🚨 AMAZON IS ENTERING THE AI CHIP MARKET — reactive:asic-gpu-market-dynamics (2026-06-19)
- [25] 🚨 $NVDA -NVIDIA CEO JENSEN HUANG: NEXT YEAR’S GROWTH WILL REMAIN RAPID 🚀🧠 — reactive:asic-gpu-market-dynamics (2026-06-14)
- [26] Jensen Huang showed up at a Marvell keynote to explain why connectivity matters (Save this). — Milk Road AI Twitter (2026-06-19)
- [27] OpenAI and Broadcom unveiled first chip Jalapeno for LLM inference. OpenAI led architecture, Broadcom handled silicon an... — reactive:asic-gpu-market-dynamics (2026-06-25)
- [28] OpenAI Unveils 'Jalapeño' Inference Chip - AIToday — reactive:asic-gpu-market-dynamics
- [29] What Is OpenAI's Jalapeno Chip? The Custom AI Inference Processor Explained | MindStudio — reactive:asic-gpu-market-dynamics
- [30] How OpenAI Built the Jalapeño Custom Inference Chip [Model Behavior] — reactive:asic-gpu-market-dynamics
- [31] Tensordyne Announces Breakthrough Inference System - LinkedIn — reactive:asic-gpu-market-dynamics
- [32] Broadcom and OpenAI unveil custom-built Jalapeño inference ... — reactive:asic-gpu-market-dynamics