The Information Machine

NVIDIA vs. Custom ASICs: GPU Dominance Persists Despite Startup Performance Claims · history

Version 9

2026-07-02 08:32 UTC · 108 items

What

NVIDIA has held or grown AI compute market share against custom ASICs in mid-2026 [1], while its NVLink Fusion ecosystem spans foundries, chip designers, and connectivity vendors [7][9]. OpenAI's Jalapeño, an LLM inference chip built with Broadcom in nine months, entered production targeting gigawatt-scale deployment with Microsoft starting in 2026, claiming better performance-per-watt and lower serving costs than current accelerators — claims without independent benchmarks [10][11][12]. Inference startup Etched drew technical attention for a cluster-scale memory (CSM) architecture; Andrej Karpathy praised its extreme low-voltage engineering for maximizing tokens-per-watt at interactive speeds [18], while SemiAnalysis noted the high SerDes count the design requires [17].

Why it matters

The central question — whether software ecosystem depth or workload-specific hardware design delivers better AI compute economics — is now being tested by multiple credible actors simultaneously: OpenAI with captive model workloads, Etched with a technically novel memory architecture that has attracted credible endorsement, and NVIDIA with compounding software optimization that cut serving costs 2.5x in under 70 days [4]. Outcomes will determine whether CUDA's software moat holds as the dominant structural advantage or whether hardware-first approaches can accumulate enough workload ownership to overcome it.

Open questions

  • Can Jalapeño's claimed performance-per-watt and serving cost advantages survive independent benchmarking, given that no external validation has been published? [10][11]

  • Does Etched's CSM architecture — shared low-latency memory pool across a scale-up domain, high SerDes count — translate to verified production advantages over HBM-based systems, or does complexity offset the architectural gains? [18][17]

  • Does the 2.5x serving cost reduction on NVIDIA's GB200 NVL72 via CUDA kernel rewrites in under 70 days represent a rate of software improvement that inference-specific ASICs cannot match on near-term timelines? [4]

  • Will Tensordyne's 13x throughput claim [15] or DeepAdapt's 82% cost reduction [16] survive independent testing, given both rest on unverified internal benchmarks?

Narrative

For roughly two years, the dominant expectation in AI compute markets was that custom ASICs would gradually absorb workloads from NVIDIA GPUs. Mid-2026 data has not confirmed that prediction: NVIDIA has held or grown its market share against ASICs [1], and CEO Jensen Huang has publicly projected continued rapid growth [2]. The analytical explanation that has gained the most traction centers on software. SemiAnalysis argues that AI chip competition is decided by software ecosystem depth, not hardware benchmarks — that 99% of chip startups publish impressive simulated numbers but fail in production because building a competitive software stack is the actual hard problem, one that NVIDIA's CUDA platform has largely solved over two decades [3]. A concrete data point supports this argument: serving costs on NVIDIA's GB200 NVL72 dropped 2.5x in under 70 days through CUDA kernel rewrites applied to the Kimi model architecture [4], illustrating software-layer compounding that hardware-first competitors cannot replicate on short timelines.

NVIDIA has structured custom silicon as a participant in its own ecosystem rather than an external competitor. NVLink Fusion allows hyperscalers and chip designers to integrate non-NVIDIA compute into NVIDIA-led clusters [5]. The ecosystem spans Marvell (hyperscaler ASICs), SiFive (RISC-V AI data center chips) [6], Samsung Foundry as a manufacturing partner [7], Astera Labs (connectivity silicon) [8], and MediaTek [9], covering design, manufacturing, and connectivity layers.

On June 24, 2026, OpenAI and Broadcom announced Jalapeño — OpenAI's first custom Intelligence Processor, designed for LLM inference and built from design to tape-out in nine months [10][11]. OpenAI claims better performance per watt than current accelerators in early testing and lower model-serving costs [12], targeting gigawatt-scale deployment with Microsoft starting in 2026 [13]. SemiAnalysis noted the nine-month development timeline requires 'making no mistakes' [14], extending its ASIC skepticism to OpenAI's claim. What differentiates Jalapeño from startup hardware claims is that OpenAI brings captive workloads and model-development depth — exactly the software-stack ownership SemiAnalysis argues most ASIC projects lack. No independent benchmarks have been published. Hardware startups Tensordyne and DeepAdapt have each published large unverified performance claims — 13x rack throughput versus NVIDIA's NVL72 GB300 on DeepSeek-R1 [15] and 82% AI operating cost reduction by shifting workloads from GPUs to CPUs [16] — both based on internal simulations, fitting the benchmark-slide pattern SemiAnalysis describes.

Inference startup Etched has attracted credible technical attention for a distinct architectural approach. Its cluster-scale memory (CSM) creates a shared low-latency memory pool across an entire scale-up domain using a proprietary interconnect, combining HBM and SRAM in a hybrid designed to achieve both high throughput and interactive inference speeds without the tradeoffs of SRAM-only designs or 3D DRAM [17]. Andrej Karpathy publicly praised the engineering approach, describing the design as operating at 'very low voltage domains' with 'cluster scale memory' and framing efficient LLM inference as the electrical inverse of power transmission — very low voltage, high current at tiny distances [18]. SemiAnalysis noted the high SerDes count the architecture requires [17], a technically neutral observation pointing to both the novelty and the complexity of the approach. No independent benchmarks of Etched's system have been published.

Timeline

  • 2026-01-15: SiFive announces integration of NVLink Fusion for next-generation RISC-V AI data center chips. [6]
  • 2026-03: Meta announces expansion of its MTIA custom silicon program to power next-generation AI workloads. [23]
  • 2026-06-14: NVIDIA CEO Jensen Huang states publicly that next year's growth will remain rapid. [2]
  • 2026-06-17: Tensordyne announces an inference rack claiming 13x rack throughput versus NVIDIA's NVL72 GB300 on DeepSeek-R1, based on internal simulations. [15][25]
  • 2026-06-18: Milk Road AI publishes analysis arguing data shows NVIDIA has held or grown market share against ASICs, contrary to two years of consensus predictions. [1]
  • 2026-06-19: SemiAnalysis states that 99% of custom ASIC projects fail and that AI chip success is determined by software depth, not hardware specifications. [3]
  • 2026-06-19: Jensen Huang appears at a Marvell keynote framing NVLink Fusion as enabling custom silicon coexistence within NVIDIA-led clusters. [5]
  • 2026-06-19: Amazon reported to be entering the AI chip market. [24]
  • 2026-06-19: DeepAdapt claims its runtime layer cuts AI operating costs by 82% and delivers 33x faster inference by shifting workloads from GPUs to CPUs, without independent validation. [16]
  • 2026-06-21: Samsung Foundry added to NVIDIA's NVLink Fusion ecosystem as a manufacturing partner for custom silicon. [7]
  • 2026-06-22: SemiAnalysis documents 2.5x serving cost reduction on GB200 NVL72 in under 70 days via CUDA kernel rewrites. [4]
  • 2026-06-24: OpenAI and Broadcom announce Jalapeño, OpenAI's first custom LLM inference chip built in nine months, claiming better performance per watt; deployment with Microsoft targets gigawatt scale starting in 2026. [10][11][20][13]
  • 2026-06-24: SemiAnalysis notes the Jalapeño nine-month development timeline requires 'making no mistakes.' [14]
  • 2026-06-30: Andrej Karpathy praises Etched's cluster-scale memory architecture for extreme low-voltage engineering targeting tokens-per-watt at interactive inference speeds; SemiAnalysis notes the design's high SerDes count. [18][17]

Perspectives

NVIDIA / Jensen Huang

Projects continued rapid growth; NVLink Fusion positions custom silicon makers as ecosystem participants rather than competitors, with NVIDIA as the interconnect backbone of AI compute.

Evolution: Consistently bullish; NVLink Fusion has expanded from a few large hyperscalers to chip designers, foundries, consumer silicon makers, and connectivity vendors.

OpenAI / Broadcom

Jalapeño is OpenAI's first custom Intelligence Processor for LLM inference, co-developed with Broadcom in nine months; early tests claim better performance per watt and lower serving costs than current accelerators, targeting gigawatt-scale deployment with Microsoft in 2026.

Evolution: First appeared June 24; secondary coverage frames the effort as NVIDIA dependency reduction; no independent benchmarks yet.

SemiAnalysis

AI chip competition is decided by software ecosystem depth; 99% of ASIC projects fail in production regardless of simulation results; Jalapeño's nine-month timeline requires 'making no mistakes'; Etched's CSM design has a very high SerDes count.

Evolution: Consistent skeptic of hardware-first ASIC narratives; June 2026 CUDA cost-reduction data sharpens the software moat argument; technical commentary now extends to Etched's architecture as well as OpenAI's timeline.

Milk Road AI

Data shows NVIDIA has held or grown market share against ASICs; NVLink Fusion makes custom silicon makers ecosystem participants rather than competitors.

Evolution: Contrarian on ASIC displacement; consistently reframes competition as coexistence via NVLink Fusion.

NVLink Fusion ecosystem (Marvell, SiFive, Samsung Foundry, MediaTek, Astera Labs)

Chip designers, foundries, and connectivity silicon makers treat NVLink Fusion as shared infrastructure, operating within NVIDIA's interconnect framework rather than against it.

Evolution: Expanded from Marvell alone to a set spanning manufacturing, consumer chip design, RISC-V AI chips, and connectivity.

Meta / Amazon (hyperscalers)

Both treating custom silicon as critical for scaling AI workloads — Meta expanding MTIA, Amazon entering the AI chip market — with captive workloads and engineering scale that differs structurally from hardware-first startups.

Evolution: Deepening commitment to custom silicon through 2026; structurally distinct from startups in ability to develop accompanying software.

Tensordyne / DeepAdapt

Both claim large performance advantages over NVIDIA-centric inference (13x throughput and 82% cost reduction respectively), based on internal simulations without independent validation.

Evolution: Fit the benchmark-slide pattern SemiAnalysis describes; neither claim has been independently verified.

Etched

Claims its CSM architecture creates a shared low-latency memory pool across a scale-up domain, enabling high throughput and interactive inference speeds while avoiding the tradeoffs of SRAM-only or 3D DRAM approaches.

Evolution: New to this thread as of June 30; drew public technical endorsement from Andrej Karpathy and architectural commentary from SemiAnalysis, but no independent benchmarks have been published.

Tensions

  • SemiAnalysis argues the CUDA software moat is real and deepening — 2.5x serving cost reduction in 70 days via kernel rewrites — while OpenAI claims Jalapeño delivers better performance per watt through workload-specific hardware design; SemiAnalysis's 'making no mistakes' comment signals skepticism of OpenAI's timeline. [4][10][14]
  • SemiAnalysis argues AI chip competition is decided by software depth and 99% of ASIC projects fail in production, while Tensordyne, DeepAdapt, and Etched publish architectural claims framed as improvements over NVIDIA hardware, none with independent validation. [3][15][16][18][17]
  • Milk Road AI and NVIDIA argue the 'NVIDIA versus everyone' frame is wrong — NVLink Fusion makes custom silicon a coexistent participant; SemiAnalysis supports NVIDIA's resilience through a different mechanism: software moat, not ecosystem embrace. [5][1][3]
  • Milk Road AI argues mid-2026 data shows NVIDIA holding or growing market share, directly contradicting two years of consensus predictions that custom silicon would erode NVIDIA's position. [1]
  • The software moat argument holds that startups cannot overcome CUDA's depth; Meta, Amazon, and OpenAI — organizations with captive workloads and engineering scale — test whether sufficient resources can develop the required software stack alongside custom hardware. [3][23][24][10]
  • Karpathy publicly endorses Etched's extreme low-voltage, cluster-scale memory architecture as technically impressive; SemiAnalysis's pointed note on the high SerDes count implicitly flags architectural complexity, but neither voice has addressed production viability directly. [18][17]

Sources

  1. [1] Everyone assumed Nvidia would get crushed by ASICs but the data says the opposite just happened and the reason why chang… — Milk Road AI Twitter (2026-06-18)
  2. [2] 🚨 $NVDA -NVIDIA CEO JENSEN HUANG: NEXT YEAR’S GROWTH WILL REMAIN RAPID 🚀🧠 — reactive:asic-gpu-market-dynamics (2026-06-14)
  3. [3] 100% of AI chip startups have slides/“simulated performance data” showing that their chip is way better, but 99% of cust… — SemiAnalysis Twitter (2026-06-19)
  4. [4] CUDA MOAT ALERT 🔥: In less than 70 days, GB200 NVL72 serving costs decreased by 2.5x through software improvements alone… — SemiAnalysis Twitter (2026-06-22)
  5. [5] Jensen Huang showed up at a Marvell keynote to explain why connectivity matters (Save this). — Milk Road AI Twitter (2026-06-19)
  6. [6] SiFive To Integrate Nvidia NVLink Fusion For Next-Gen AI Data Centers — reactive:asic-gpu-market-dynamics
  7. [7] NVIDIA Adds Samsung Foundry to NVLink Fusion Ecosystem for Custom Silicon Manufacturing : r/hardware — reactive:asic-gpu-market-dynamics
  8. [8] Why Connectivity is the New Frontier of AI Infrastructure—and What NVLink Fusion Means for the Future — reactive:asic-gpu-market-dynamics
  9. [9] MediaTek | NVLink Fusion | Custom AI ASIC Innovation — reactive:asic-gpu-market-dynamics
  10. [10] OpenAI and Broadcom unveil LLM-optimized inference chip — OpenAI Blog (2026-06-24)
  11. [11] OpenAI and Broadcom announce chip designed for LLM inference at scale — Ars Technica AI (2026-06-24)
  12. [12] 🟡 Hollywood's AI trick work — Semafor Technology (2026-06-26)
  13. [13] OpenAI and Broadcom unveiled first chip Jalapeno for LLM inference. OpenAI led architecture, Broadcom handled silicon an... — reactive:asic-gpu-market-dynamics (2026-06-25)
  14. [14] Chat develop a chip from initial design to tape out in 9 months, make no mistakes. https://t.co/DaNCCPJebw — SemiAnalysis Twitter (2026-06-24)
  15. [15] Quite a massive inferencing rack breakthrough from @TensordyneInc . — Rohan Paul Twitter (2026-06-17)
  16. [16] DeepAdapt has launched a runtime intelligence layer that cuts AI operating costs by up to 82% and 33X faster inference b… — Rohan Paul Twitter (2026-06-19)
  17. [17] etched cluster-scale memory has so many SerDes https://t.co/V6yTAtsYf9 https://t.co/OZHcQshvUj — SemiAnalysis Twitter (2026-06-30)
  18. [18] Congrats!! I was impressed to learn about some of the engineering wizardry (e.g. *very* low voltage domains, cluster sca… — Andrej Karpathy Twitter (2026-06-30)
  19. [19] Build Semi-Custom AI Infrastructure | NVIDIA NVLink Fusion — reactive:asic-gpu-market-dynamics
  20. [20] Broadcom and OpenAI unveil custom-built Jalapeño inference ... — reactive:asic-gpu-market-dynamics
  21. [21] OpenAI unveils "Jalapeno" custom AI chip to reduce Nvidia reliance, promising superior inference efficiency and 2026 dep... — reactive:asic-gpu-market-dynamics (2026-06-26)
  22. [22] In 2016, Marvell's largest design win was a Wi-Fi chip in the Barbie Dream House (Save this). — Milk Road AI Twitter (2026-06-20)
  23. [23] Expanding Meta's Custom Silicon to Power Our AI Workloads — reactive:asic-gpu-market-dynamics
  24. [24] 🚨 AMAZON IS ENTERING THE AI CHIP MARKET — reactive:asic-gpu-market-dynamics (2026-06-19)
  25. [25] Tensordyne Announces Breakthrough Inference System - LinkedIn — reactive:asic-gpu-market-dynamics