The Information Machine

NVIDIA vs. Custom ASICs: GPU Dominance Persists Despite Startup Performance Claims · history

Version 5

2026-06-25 18:28 UTC · 59 items

What

NVIDIA has held or grown AI compute market share against custom ASICs as of mid-2026 [1], while its NVLink Fusion ecosystem now encompasses chip designers, foundries, and connectivity silicon makers including Samsung Foundry [7] and MediaTek [9]. On June 24, 2026, OpenAI and Broadcom announced 'Jalapeño' — OpenAI's first custom LLM inference chip, developed in nine months, claiming better performance per watt than current accelerators [10]. OpenAI is the first major AI model operator to enter custom silicon for inference, bringing workload-specific depth that differs structurally from hardware startups — but the chip has not been independently benchmarked.

Why it matters

OpenAI's entry into custom silicon is different in kind from startup ASIC claims: OpenAI owns both the models and the software stack, which is exactly the gap SemiAnalysis argues strands most ASIC projects. Whether the software moat argument holds against an operator that already controls the workloads is now an open empirical question rather than a theoretical one.

Open questions

  • Can OpenAI's workload-specific design knowledge translate into production advantages at scale, or will Jalapeño's early performance claims [10] follow the benchmark-slide pattern SemiAnalysis describes for hardware startups [3]?

  • Does the 2.5x serving cost reduction through CUDA kernel rewrites in under 70 days [4] represent a rate of software improvement that inference-specific ASICs cannot close over time?

  • Will Tensordyne's 13x throughput claim [13] or DeepAdapt's 82% cost reduction [14] survive independent testing, given that both rest on unverified internal benchmarks?

  • Does NVLink Fusion's expansion to Samsung Foundry and MediaTek [7][9] mean NVIDIA captures value from the manufacturing and consumer chip design layers, not just hyperscaler ASICs?

Narrative

For roughly two years, the dominant expectation in AI compute markets was that custom ASICs would gradually absorb workloads from NVIDIA GPUs. Mid-2026 data has not confirmed that prediction: NVIDIA has held or grown its market share against ASICs [1], and Jensen Huang has publicly projected continued rapid growth [2].

The analytical explanation that has gained the most traction centers on software. SemiAnalysis argues that AI chip competition is decided by software ecosystem depth, not hardware benchmarks: 99% of chip startups publish impressive simulated numbers but fail in production because building a competitive software stack alongside hardware is the actual hard problem — one that NVIDIA's CUDA platform has largely solved over two decades [3]. A concrete data point: GB200 NVL72 serving costs dropped 2.5x in under 70 days through CUDA kernel rewrites alone applied to the Kimi model architecture [4], illustrating software-layer compounding that hardware-first competitors cannot replicate on near-term timelines.

NVIDIA has also structured custom silicon as a participant in its own ecosystem rather than an external competitor. NVLink Fusion allows hyperscalers and chip designers to integrate non-NVIDIA compute into NVIDIA-led clusters [5]. The ecosystem spans Marvell (hyperscaler ASICs), SiFive (RISC-V AI data center chips) [6], Samsung Foundry as a manufacturing partner [7], Astera Labs (connectivity silicon) [8], and MediaTek [9] — now covering design, manufacturing, and connectivity layers.

On June 24, 2026, OpenAI and Broadcom announced 'Jalapeño' — OpenAI's first custom Intelligence Processor, designed for LLM inference and built from design to tape-out in nine months, which OpenAI and Broadcom claim is the fastest such development cycle ever achieved in high-performance advanced semiconductors [10][11]. OpenAI claims better performance per watt than current accelerators in early testing, used its own AI models to accelerate chip design, and targets gigawatt-scale deployment with Microsoft starting in 2026. SemiAnalysis noted the nine-month timeline with a pointed observation that such a process requires 'making no mistakes' [12]. Unlike the unverified claims from Tensordyne (13x throughput [13]) and DeepAdapt (82% cost reduction [14]), OpenAI brings captive workloads and model-development depth to its hardware effort — but Jalapeño has not yet been independently benchmarked.

Timeline

  • 2026-01-15: SiFive announces integration of NVLink Fusion for next-generation RISC-V AI data center chips. [6]
  • 2026-03: Meta announces expansion of its MTIA custom silicon program to power next-generation AI workloads. [19]
  • 2026-06-14: NVIDIA CEO Jensen Huang states publicly that next year's growth will remain rapid. [2]
  • 2026-06-15: Analysis published on China's growing supply chain independence in AI hardware, covering compound semiconductors and Huawei's architecture work. [21]
  • 2026-06-17: Tensordyne announces an inference rack claiming 13x rack throughput versus NVIDIA's NVL72 GB300 on DeepSeek-R1, based on internal simulations. [13][22]
  • 2026-06-18: Milk Road AI publishes analysis arguing data shows NVIDIA has held or grown market share against ASICs, contrary to two years of consensus predictions. [1]
  • 2026-06-19: Jensen Huang appears at a Marvell keynote framing NVLink Fusion as enabling custom silicon coexistence within NVIDIA-led clusters. [5]
  • 2026-06-19: Amazon reported to be entering the AI chip market. [20]
  • 2026-06-19: SemiAnalysis states that 99% of custom ASIC projects fail and that AI chip success is determined by software depth, not hardware specifications. [3]
  • 2026-06-19: DeepAdapt claims its runtime layer cuts AI operating costs by 82% and delivers 33x faster inference by shifting workloads from GPUs to CPUs, without independent validation. [14]
  • 2026-06-21: Samsung Foundry added to NVIDIA's NVLink Fusion ecosystem as a manufacturing partner for custom silicon. [7]
  • 2026-06-22: SemiAnalysis documents 2.5x serving cost reduction on GB200 NVL72 in under 70 days via CUDA kernel rewrites. [4]
  • 2026-06-24: OpenAI and Broadcom announce 'Jalapeño,' OpenAI's first custom LLM inference chip developed in nine months, claiming better performance per watt than current accelerators. [10][11]

Perspectives

NVIDIA / Jensen Huang

Projects continued rapid growth; NVLink Fusion positions custom silicon makers as ecosystem participants rather than competitors, with NVIDIA as the interconnect backbone of AI compute.

Evolution: Consistently bullish on growth; NVLink Fusion has expanded from a few large hyperscalers to a broad set including chip designers, foundries, and consumer chip makers such as MediaTek.

OpenAI / Broadcom

Jalapeño is OpenAI's first custom Intelligence Processor, designed from scratch for LLM inference and co-developed with Broadcom in nine months; early tests show better performance per watt than current accelerators, targeting gigawatt-scale deployment with Microsoft in 2026.

Evolution: New voice in this thread; first appearance with the June 24 announcement.

SemiAnalysis

AI chip competition is decided by software ecosystem depth; 99% of ASIC projects fail in production regardless of simulation results. The 2.5x serving cost reduction via CUDA rewrites in 70 days is direct production evidence. Noted OpenAI's 9-month chip development timeline with the observation that such a process requires 'making no mistakes.'

Evolution: Consistent skeptic of hardware-first ASIC narratives; June 2026 CUDA cost-reduction data sharpens the software moat argument, while the Jalapeño comment implies scrutiny of the timeline claim.

Milk Road AI

Data shows NVIDIA has held or grown market share against ASICs; NVLink Fusion makes custom silicon makers ecosystem participants rather than competitors.

Evolution: Contrarian on ASIC displacement; consistently reframes competition as coexistence via NVLink Fusion.

NVLink Fusion ecosystem participants (Marvell, SiFive, Samsung Foundry, MediaTek, Astera Labs)

Chip designers, foundries, and connectivity silicon makers treat NVLink Fusion as shared infrastructure, operating within NVIDIA's interconnect framework rather than against it.

Evolution: Expanded from Marvell alone to a set now spanning manufacturing (Samsung), consumer chip design (MediaTek), RISC-V AI chips (SiFive), and connectivity (Astera Labs).

Meta / Amazon (hyperscalers)

Both treating custom silicon as critical for scaling AI workloads — Meta expanding its MTIA program, Amazon entering the AI chip market — with captive workloads and engineering scale that potentially differs from hardware-first startups.

Evolution: Deepening commitment to custom silicon through 2026; structurally distinct from startups in ability to develop supporting software.

Tensordyne / DeepAdapt

Both claim large performance advantages over NVIDIA-centric inference (13x throughput and 82% cost reduction respectively), based on internal simulations without independent validation.

Evolution: Fit the benchmark-slide pattern SemiAnalysis describes; neither claim has been independently verified.

Sky Rain

China's positions in compound semiconductors and Huawei's architecture work are building a supply chain track that operates outside the NVIDIA-dominated ecosystem.

Evolution: Consistent focus on China supply chain independence as a structurally distinct competitive thread.

Tensions

  • SemiAnalysis argues the CUDA software moat is real and deepening — 2.5x serving cost reduction in 70 days via kernel rewrites — while OpenAI claims Jalapeño delivers better performance per watt through workload-specific hardware design; SemiAnalysis's comment about 'making no mistakes' signals skepticism of the timeline. [4][10][12]
  • SemiAnalysis argues AI chip competition is decided by software depth and 99% of ASIC projects fail in production, while Tensordyne and DeepAdapt publish large unverified numbers framed as breakthroughs over NVIDIA hardware. [3][13][14]
  • Milk Road AI and NVIDIA argue the 'NVIDIA versus everyone' frame is wrong — NVLink Fusion makes custom silicon a coexistent participant; SemiAnalysis supports NVIDIA's resilience but through a different mechanism (software moat, not ecosystem embrace). [5][1][3]
  • Milk Road AI argues mid-2026 data shows NVIDIA holding or growing market share, directly contradicting two years of consensus predictions that custom silicon would erode NVIDIA's position. [1]
  • The software moat argument holds that startups cannot overcome CUDA's depth; Meta, Amazon, and now OpenAI — organizations with captive workloads and engineering scale — test whether sufficient resources can develop the required software stack alongside custom hardware. [3][19][20][10]
  • Tensordyne's 13x throughput claim and DeepAdapt's 82% cost reduction both remain unverified by any third party; the gap between simulation-based numbers and any independent result is unresolved. [13][14]

Sources

  1. [1] Everyone assumed Nvidia would get crushed by ASICs but the data says the opposite just happened and the reason why chang… — Milk Road AI Twitter (2026-06-18)
  2. [2] 🚨 $NVDA -NVIDIA CEO JENSEN HUANG: NEXT YEAR’S GROWTH WILL REMAIN RAPID 🚀🧠 — reactive:asic-gpu-market-dynamics (2026-06-14)
  3. [3] 100% of AI chip startups have slides/“simulated performance data” showing that their chip is way better, but 99% of cust… — SemiAnalysis Twitter (2026-06-19)
  4. [4] CUDA MOAT ALERT 🔥: In less than 70 days, GB200 NVL72 serving costs decreased by 2.5x through software improvements alone… — SemiAnalysis Twitter (2026-06-22)
  5. [5] Jensen Huang showed up at a Marvell keynote to explain why connectivity matters (Save this). — Milk Road AI Twitter (2026-06-19)
  6. [6] SiFive To Integrate Nvidia NVLink Fusion For Next-Gen AI Data Centers — reactive:asic-gpu-market-dynamics
  7. [7] NVIDIA Adds Samsung Foundry to NVLink Fusion Ecosystem for Custom Silicon Manufacturing : r/hardware — reactive:asic-gpu-market-dynamics
  8. [8] Why Connectivity is the New Frontier of AI Infrastructure—and What NVLink Fusion Means for the Future — reactive:asic-gpu-market-dynamics
  9. [9] MediaTek | NVLink Fusion | Custom AI ASIC Innovation — reactive:asic-gpu-market-dynamics
  10. [10] OpenAI and Broadcom unveil LLM-optimized inference chip — OpenAI Blog (2026-06-24)
  11. [11] OpenAI and Broadcom announce chip designed for LLM inference at scale — Ars Technica AI (2026-06-24)
  12. [12] Chat develop a chip from initial design to tape out in 9 months, make no mistakes. https://t.co/DaNCCPJebw — SemiAnalysis Twitter (2026-06-24)
  13. [13] Quite a massive inferencing rack breakthrough from @TensordyneInc . — Rohan Paul Twitter (2026-06-17)
  14. [14] DeepAdapt has launched a runtime intelligence layer that cuts AI operating costs by up to 82% and 33X faster inference b… — Rohan Paul Twitter (2026-06-19)
  15. [15] Build Semi-Custom AI Infrastructure | NVIDIA NVLink Fusion — reactive:asic-gpu-market-dynamics
  16. [16] OpenAI rolls out its 1st chip through a Broadcom tie-up as part of its “build the full stack” push. — Rohan Paul Twitter (2026-06-24)
  17. [17] 😺 ChatGPT's secret advantage — The Neuron (2026-06-25)
  18. [18] In 2016, Marvell's largest design win was a Wi-Fi chip in the Barbie Dream House (Save this). — Milk Road AI Twitter (2026-06-20)
  19. [19] Expanding Meta's Custom Silicon to Power Our AI Workloads — reactive:asic-gpu-market-dynamics
  20. [20] 🚨 AMAZON IS ENTERING THE AI CHIP MARKET — reactive:asic-gpu-market-dynamics (2026-06-19)
  21. [21] China’s Quiet Dominance in the AI Supply Chain: InP, Semiconductor Independence, and Huawei’s Logic Folding Wildcard — reactive:asic-gpu-market-dynamics (2026-06-15)
  22. [22] Tensordyne Announces Breakthrough Inference System - LinkedIn — reactive:asic-gpu-market-dynamics