NVIDIA vs. Custom ASICs: GPU Dominance Persists Despite Startup Performance Claims · history
Version 4
2026-06-24 02:30 UTC · 53 items
What
NVIDIA has held or grown AI compute market share against custom ASICs as of mid-2026 [1], while its NVLink Fusion ecosystem now spans chip designers (Marvell, SiFive), manufacturing (Samsung Foundry), and connectivity silicon (Astera Labs) [6][7][8]. A concrete software-layer data point arrived in June 2026: GB200 NVL72 serving costs fell 2.5x in under 70 days through CUDA kernel rewrites alone [4], illustrating why startup benchmark claims rest on fundamentally different ground than production software stacks. Startup performance claims — Tensordyne's 13x throughput [9] and DeepAdapt's 82% cost reduction [10] — remain unverified by any third party.
Why it matters
NVIDIA's software stack compounds independently of hardware releases, continuously widening the gap between simulation-based startup benchmarks and production performance. The NVLink Fusion strategy converts potential competitors into ecosystem participants, capturing value from custom silicon growth rather than losing ground to it.
Open questions
Does the 2.5x serving cost reduction through CUDA kernel rewrites in under 70 days represent a structural rate of software improvement, or a one-time optimization that competitors could match? [4]
Will Tensordyne's 13x throughput claim or DeepAdapt's 82% cost reduction survive independent testing, given that both rest on unverified internal benchmarks? [9][10]
Will Meta and Amazon's hyperscaler custom silicon programs remain within NVIDIA's NVLink Fusion ecosystem or develop independent software stacks at sufficient depth? [11][12]
Does Samsung Foundry's addition to NVLink Fusion mean NVIDIA now captures value from the manufacturing layer, not just chip design — and how durable is that position? [6]
Narrative
For roughly two years, the dominant expectation in AI compute markets was that custom ASICs — built by hyperscalers and startups alike — would gradually absorb workloads from NVIDIA GPUs. Mid-2026 data has not confirmed that prediction: NVIDIA has held or grown its market share against ASICs [1], and Jensen Huang has publicly projected continued rapid growth into the following year [2].
The analytical explanation that has gained the most traction centers on software. SemiAnalysis argues that AI chip competition is decided by software ecosystem depth: 100% of chip startups publish slides with impressive simulated benchmark numbers, but 99% fail in production because building a competitive software stack alongside hardware is the actual hard problem — one that NVIDIA's CUDA platform, accumulated over roughly two decades, has largely solved [3]. A concrete data point arrived in June 2026: GB200 NVL72 serving costs dropped 2.5x in under 70 days through CUDA kernel rewrites alone, applied to the Kimi model architecture used in xAI's Cursor Composer 2.5 [4]. That kind of software-layer compounding is structurally difficult for a hardware-first competitor to replicate on any near-term timeline.
NVIDIA has also moved to make custom silicon a structural participant in its own ecosystem rather than an external competitor. NVLink Fusion allows hyperscalers and chip designers to integrate non-NVIDIA compute silicon into NVIDIA-led clusters [5]. Jensen Huang appeared at a Marvell keynote to explain this coexistence model; the ecosystem has since expanded to Samsung Foundry as a manufacturing partner [6], SiFive for RISC-V-based AI data center chips [7], and Astera Labs — which makes connectivity silicon for AI clusters — framing NVLink Fusion as making connectivity the central layer of AI infrastructure [8]. The practical effect is that custom silicon makers, whether designing chips or manufacturing them, increasingly operate within NVIDIA's interconnect framework rather than outside it.
The startup benchmark cycle continues alongside these structural developments. Tensordyne published a claim of 13x rack throughput versus NVIDIA's NVL72 GB300 on DeepSeek-R1 workloads, based on internal simulations with no independent validation [9]. DeepAdapt claims its runtime layer cuts AI operating costs by 82% and delivers 33x faster inference by shifting workloads from GPUs to CPUs [10] — also unverified. Hyperscalers are a structurally distinct case: Meta has expanded its MTIA custom silicon program [11] and Amazon is entering the AI chip market [12]. Unlike independent startups, these organizations have the engineering scale and captive workloads to potentially develop supporting software stacks alongside custom hardware. China's domestic AI supply chain — compound semiconductors and Huawei's architecture work [13] — operates outside this ecosystem frame entirely as a separate competitive track.
Timeline
- 2026-01-15: SiFive announces integration of NVLink Fusion for next-generation RISC-V AI data center chips. [7]
- 2026-03: Meta announces expansion of its MTIA custom silicon program to power next-generation AI workloads. [11]
- 2026-06-14: NVIDIA CEO Jensen Huang states publicly that next year's growth will remain rapid. [2]
- 2026-06-15: Analysis published on China's growing supply chain independence in AI hardware, covering compound semiconductors and Huawei's architecture work. [13]
- 2026-06-17: Tensordyne announces an inference rack claiming 13x rack throughput versus NVIDIA's NVL72 GB300 on DeepSeek-R1, based on internal simulations. [9][16]
- 2026-06-18: Milk Road AI publishes analysis arguing data shows NVIDIA has held or grown market share against ASICs, contrary to two years of consensus predictions. [1]
- 2026-06-19: Jensen Huang appears at a Marvell keynote to explain connectivity's role in AI systems, framing NVLink Fusion as enabling custom silicon coexistence within NVIDIA-led clusters. [5]
- 2026-06-19: Amazon reported to be entering the AI chip market. [12]
- 2026-06-19: SemiAnalysis states that 99% of custom ASIC projects fail and that AI chip success is determined by software depth, not hardware specifications. [3]
- 2026-06-19: DeepAdapt claims its runtime layer cuts AI operating costs by 82% and delivers 33x faster inference by shifting workloads from GPUs to CPUs, without independent validation. [10]
- 2026-06-20: Milk Road AI highlights Marvell's transformation from consumer chip maker to major AI semiconductor supplier via $36 billion in acquisitions over a decade. [15]
- 2026-06-21: Samsung Foundry added to NVIDIA's NVLink Fusion ecosystem as a manufacturing partner for custom silicon. [6]
- 2026-06-21: Astera Labs publishes analysis framing NVLink Fusion as making connectivity the central layer of AI infrastructure. [8]
- 2026-06-22: SemiAnalysis documents 2.5x serving cost reduction on GB200 NVL72 in under 70 days via CUDA kernel rewrites, citing it as evidence of a deepening software moat. [4]
Perspectives
NVIDIA / Jensen Huang
Projects continued rapid growth; NVLink Fusion positions custom silicon makers as ecosystem participants rather than competitors, reframing NVIDIA's role as the interconnect backbone of AI compute.
Evolution: Consistently bullish on growth; the NVLink Fusion strategy has expanded from enabling a few large players to encompassing a broad set of chip designers and foundries.
SemiAnalysis
AI chip competition is decided by software ecosystem depth, not hardware benchmarks; 99% of ASIC projects fail in production regardless of simulation results. The 2.5x serving cost reduction via CUDA rewrites in 70 days is cited as direct evidence of this moat deepening.
Evolution: Consistent skeptic of hardware-first ASIC narratives; the June 2026 CUDA cost-reduction data point sharpens the software moat argument with a concrete, dated example.
Milk Road AI
Data shows NVIDIA has held or grown market share against ASICs; the 'NVIDIA versus everyone' competitive frame is wrong and NVLink Fusion makes custom silicon makers ecosystem participants.
Evolution: Contrarian on ASIC displacement; consistently reframes competition as coexistence via NVLink Fusion.
NVLink Fusion ecosystem participants (Marvell, SiFive, Samsung Foundry)
Marvell builds hyperscaler ASICs, SiFive integrates NVLink Fusion for RISC-V AI data center chips, and Samsung Foundry manufactures within the ecosystem — all treating NVLink Fusion as shared infrastructure rather than competing against NVIDIA.
Evolution: Expanded from Marvell alone to a growing set of chip designers and a major foundry, showing the ecosystem now spans manufacturing as well as design.
Astera Labs
Frames NVLink Fusion as establishing connectivity as the central layer of AI infrastructure — a position that aligns with their own connectivity silicon business.
Evolution: Self-interested framing, but directionally consistent with evidence from other sources; consistent since entering the thread.
Meta
Treating custom silicon (MTIA) as critical to scaling next-generation AI; expanding internal chip program as a supplement or partial replacement for external GPU procurement.
Evolution: Deepening commitment to custom silicon through 2026.
Tensordyne / DeepAdapt
Both claim large performance advantages over NVIDIA-centric inference (13x throughput and 82% cost reduction respectively), based on internal simulations without independent validation.
Evolution: Fit the benchmark-slide pattern SemiAnalysis describes; neither claim has been independently verified.
Sky Rain
China's positions in compound semiconductors and Huawei's architecture work are building a supply chain track that operates outside the NVIDIA-dominated ecosystem.
Evolution: Consistent focus on China supply chain independence as a structurally distinct competitive thread.
Tensions
- SemiAnalysis argues the CUDA software moat is real and deepening — 2.5x serving cost reduction in 70 days via kernel rewrites — while Tensordyne and DeepAdapt publish large unverified numbers framed as breakthroughs over NVIDIA hardware. [4][9][10]
- Milk Road AI and NVIDIA argue the 'NVIDIA versus everyone' frame is wrong — NVLink Fusion makes custom silicon a coexistent participant; SemiAnalysis supports NVIDIA's resilience but through a different mechanism (software moat, not ecosystem embrace). [5][1][3]
- Milk Road AI argues mid-2026 data shows NVIDIA holding or growing market share, directly contradicting the two-year consensus prediction that custom silicon would erode NVIDIA's position. [1]
- The software moat argument holds that startups cannot overcome CUDA's depth; Meta and Amazon's hyperscaler buildout — with captive workloads and engineering scale — tests whether organizations with sufficient resources can develop the required software stack alongside custom hardware. [3][11][12]
- Tensordyne's 13x throughput claim and DeepAdapt's 82% cost reduction both remain unverified by any third party; the gap between simulation-based numbers and any independent result is unresolved. [9][10]
Sources
- [1] Everyone assumed Nvidia would get crushed by ASICs but the data says the opposite just happened and the reason why chang… — Milk Road AI Twitter (2026-06-18)
- [2] 🚨 $NVDA -NVIDIA CEO JENSEN HUANG: NEXT YEAR’S GROWTH WILL REMAIN RAPID 🚀🧠 — reactive:asic-gpu-market-dynamics (2026-06-14)
- [3] 100% of AI chip startups have slides/“simulated performance data” showing that their chip is way better, but 99% of cust… — SemiAnalysis Twitter (2026-06-19)
- [4] CUDA MOAT ALERT 🔥: In less than 70 days, GB200 NVL72 serving costs decreased by 2.5x through software improvements alone… — SemiAnalysis Twitter (2026-06-22)
- [5] Jensen Huang showed up at a Marvell keynote to explain why connectivity matters (Save this). — Milk Road AI Twitter (2026-06-19)
- [6] NVIDIA Adds Samsung Foundry to NVLink Fusion Ecosystem for Custom Silicon Manufacturing : r/hardware — reactive:asic-gpu-market-dynamics
- [7] SiFive To Integrate Nvidia NVLink Fusion For Next-Gen AI Data Centers — reactive:asic-gpu-market-dynamics
- [8] Why Connectivity is the New Frontier of AI Infrastructure—and What NVLink Fusion Means for the Future — reactive:asic-gpu-market-dynamics
- [9] Quite a massive inferencing rack breakthrough from @TensordyneInc . — Rohan Paul Twitter (2026-06-17)
- [10] DeepAdapt has launched a runtime intelligence layer that cuts AI operating costs by up to 82% and 33X faster inference b… — Rohan Paul Twitter (2026-06-19)
- [11] Expanding Meta's Custom Silicon to Power Our AI Workloads — reactive:asic-gpu-market-dynamics
- [12] 🚨 AMAZON IS ENTERING THE AI CHIP MARKET — reactive:asic-gpu-market-dynamics (2026-06-19)
- [13] China’s Quiet Dominance in the AI Supply Chain: InP, Semiconductor Independence, and Huawei’s Logic Folding Wildcard — reactive:asic-gpu-market-dynamics (2026-06-15)
- [14] Build Semi-Custom AI Infrastructure | NVIDIA NVLink Fusion — reactive:asic-gpu-market-dynamics
- [15] In 2016, Marvell's largest design win was a Wi-Fi chip in the Barbie Dream House (Save this). — Milk Road AI Twitter (2026-06-20)
- [16] Tensordyne Announces Breakthrough Inference System - LinkedIn — reactive:asic-gpu-market-dynamics