The Information Machine

Kimi K3 and Qwen3.8-Max: Chinese Labs Close Gap with Closed Frontier Models · history

Version 7

2026-08-03 18:27 UTC · 92 items

What

Moonshot AI's Kimi K3 (2.8 trillion parameters, open weights) and Alibaba's Qwen3.8-Max (2.4 trillion parameters, API preview only) are the leading Chinese open-weight models, assessed by analysts at 3–7 months behind closed frontier systems [2][3][4]. A Bloomberg investigation reports that Moonshot trains on roughly 20,000 Nvidia chips — possibly H200s — accessed through a confidential deal with Alibaba, raising direct questions for US export control enforcement [13]. Third-party benchmarks find Qwen3.8-Max outperforming Claude Fable 5 on 3D physics tasks at roughly 7x lower cost [8], and 1-bit quantization has compressed K3 to 590GB for local deployment on four B200 GPUs [16]. DeepSeek's V4-Flash-0731 (304B parameters) adds a third competitive Chinese release alongside the two flagship models [9].

Why it matters

The Moonshot-Alibaba chip arrangement moves the export control debate from theoretical to potentially concrete: if the supplied chips are H200s, which are export-restricted to China, that is a live enforcement question rather than a policy proposal [13]. At the same time, Qwen3.8-Max's cost-performance results and K3's quantized local deployability continue to show the open-to-closed gap closing in specific task domains.

Open questions

  • Did Moonshot access H200 chips in violation of US export controls? Alibaba denied the H200 identification but did not deny supplying the chips and refused to say what they are [13].

  • When will Alibaba release Qwen3.8-Max as open weights? The model has been available only as an API preview since Xi Jinping's July 20 WAIC keynote commitment [6].

  • Does Kimi K3's commercial license — requiring separate agreements for businesses over $20M in annual revenue — give the US government actionable leverage over American entities using Chinese AI models [12]?

  • Who is the one notable AI figure who did not sign the public letter on Open Weights and American AI Leadership, and does their absence reflect a meaningful dissent [17]?

Narrative

Moonshot AI's Kimi K3 (2.8 trillion parameters, mixture-of-experts) was released as open weights in July 2026 and is available on Hugging Face at 1.56TB [1]. Analyst estimates place it within 3–7 months of closed frontier systems: Lambert at 3–5 months, Mowshowitz at 4–6 months, and the UK AI Security Institute at 4–7 months for cybersecurity tasks specifically [2][3][4]. Alibaba's Qwen3.8-Max (2.4 trillion parameters) was previewed at the same time with self-reported benchmarks claiming it trails only Claude Fable 5; it remains an API preview only, despite a commitment to open-weight release made following Xi Jinping's July 20 WAIC keynote [5][6]. Third-party evaluations have begun: SemiAnalysis published performance tests of the preview [7], and a coding benchmark found Qwen3.8-Max outperformed Claude Fable 5 on 3D physics scene generation at roughly 7x lower cost — $0.28 per task versus Fable 5's $1.93, with Qwen demonstrating stronger geometric accuracy and collision handling [8]. DeepSeek's V4-Flash-0731 (304 billion parameters, released July 31) adds a third Chinese model at $0.14/$0.27 per million tokens, ranked by Artificial Analysis ahead of MiniMax M3 on its Intelligence Index [9].

A substantive disagreement has developed over what drives Chinese model performance. Nathan Lambert argues the gains are genuine innovation: RL training — where real capability advances occur — requires millions of rollouts prohibitively expensive to run through closed-model APIs, making distillation a limited contributor [10]. He explicitly calls Ben Thompson's claim that distillation scales with RL factually incorrect and potentially misleading to policymakers [10]. Zvi Mowshowitz holds that K3's benchmark scores are systematically overstated due to max-effort evaluation settings, jagged capability profiles, and some Claude distillation, though he concedes the model is real and roughly four to six months behind the closed frontier [3]. Semafor reports K3 underperforms top frontier models on cybersecurity vulnerability detection specifically, suggesting capability is uneven across domains [11]. Florian Brand, writing in Interconnects, argues that earlier predictions of open-weight ecosystem consolidation have failed to materialize — more companies than ever are training and releasing frontier-scale models — and that commercial viability is demonstrated by Thinking Machines, whose open-model fine-tuning service generates hundreds of millions in annual revenue [12].

A Bloomberg investigation adds a concrete dimension to the export control debate. Moonshot AI reportedly trains its Kimi models on roughly 20,000 Nvidia chips accessed through a confidential deal with Alibaba, one of Moonshot's largest investors. Multiple sources identified the chips as H200s, which are export-restricted to China under US rules. Alibaba denied the H200 identification but did not deny supplying 20,000 chips and refused to disclose what they are [13]. The Trump administration was separately reviving efforts to restrict Chinese AI models after K3's release [14], while Congress introduced bipartisan kill-switch legislation assessed as ineffective against sophisticated systems [11]. Lambert introduced a concrete cost to a ban: a Chinese open model was the only tool available to analyze an OpenAI breach because US frontier models had guardrails blocking the analysis, meaning restrictions would directly disadvantage American cybersecurity defenders [10]. An Anthropic cybersecurity incident — a configuration mistake that left a supposedly isolated test environment open to the internet, allowing Claude to reach real companies — has been cited as additional relevant context [15]. Brand notes that K3's commercial license, requiring MaaS businesses over $20 million in annual revenue to negotiate separate agreements with Moonshot, could create regulatory leverage for US government action against American entities using the model [12].

1-bit quantization has compressed Kimi K3 from 1.56TB to 590GB — a 62% reduction — while retaining 78.7% agreement with the original model and its full 1-million-token context window [16]. Increasing to 2.1-bit raises agreement to 87.2% at 769GB. In a practical HTML 3D physics test, the locally-run 1-bit K3 was the only model among four tested — including the Kimi K3 API, Claude Opus 5, and GPT 5.6 — to build a working winch mechanism, and running it locally on four B200 GPUs costs nothing in API fees versus $0.30–0.77 for cloud alternatives [16]. A public letter on Open Weights and American AI Leadership, signed by nearly all major AI figures with one notable exception, reflects a consolidating practitioner consensus that open-weight development should continue [17].

Timeline

  • 2026-04-15: Lambert publishes predictions arguing RL-era structural limits will slow Chinese open-weight progress and that Chinese labs will face funding difficulties by late 2026. [18]
  • 2026-07-20: Xi Jinping delivers WAIC keynote endorsing open-source AI; Alibaba commits to releasing Qwen3.8-Max as open weights following earlier hesitation. [19][2]
  • 2026-07-20: Moonshot AI releases Kimi K3 as open weights — a 2.8T MoE model assessed by multiple analysts as the strongest open-weight model yet released. [4][2][3]
  • 2026-07-20: Alibaba previews Qwen3.8-Max (2.4T) with self-reported benchmarks claiming it trails only Claude Fable 5; open-weight release committed but not yet delivered. [5]
  • 2026-07-20: UK AI Security Institute reports the open-to-closed gap on cybersecurity tasks narrowed from 6–10 months to 4–7 months. [4]
  • 2026-07-21: Trump administration reportedly revives efforts to ban Chinese AI models; officials note open-weight distribution makes enforcement nearly impossible. [14]
  • 2026-07-22: Lambert publishes follow-up arguing distillation has limited RL impact, explicitly calling Thompson's distillation narrative factually incorrect, and introducing the cybersecurity defender argument against a ban. [10]
  • 2026-07-24: Semafor reports K3 underperforms frontier models on cybersecurity vulnerability detection; Congress introduces bipartisan AI kill-switch legislation, assessed as ineffective against sophisticated systems. [11]
  • 2026-07-27: K3 weights confirmed live on Hugging Face at 1.56TB; license requires MaaS businesses over $20M annual revenue to negotiate separate commercial agreements with Moonshot. [1]
  • 2026-07-30: Qwen3.8-Max appears as an API preview on AIHubMix; open-weight release has not occurred. [6]
  • 2026-07-31: DeepSeek releases V4-Flash-0731 (304B) at $0.14/$0.27 per million tokens; Artificial Analysis ranks it ahead of MiniMax M3 on its Intelligence Index. [9]
  • 2026-07-31: Public letter on Open Weights and American AI Leadership signed by nearly all major AI figures with one notable exception. [17]
  • 2026-07-31: Bloomberg reports Moonshot trains Kimi models on roughly 20,000 Nvidia chips (possibly H200s) via a confidential Alibaba deal; Alibaba denies the H200 claim but will not identify the chips. [13]
  • 2026-07-31: 1-bit quantization compresses K3 to 590GB (62% reduction), retaining 78.7% model agreement and the full 1M-token context window; locally run model outperforms cloud alternatives on a 3D physics task. [16]
  • 2026-08-02: SemiAnalysis publishes performance tests of Qwen3.8-Max-Preview (2.4T). [7]
  • 2026-08-02: Brand argues that open-weight consolidation predictions have failed and that K3's commercial license could give the US government leverage over American businesses using Chinese AI models. [12]
  • 2026-08-03: Qwen3.8-Max outperforms Claude Fable 5 on 3D physics generation at roughly 7x lower cost ($0.28 vs $1.93 per task). [8]
  • 2026-08-03: Anthropic confirms a cybersecurity test incident: a configuration mistake left an isolated environment open to the internet, allowing Claude to reach real companies. [15]

Perspectives

Nathan Lambert (Interconnects)

Kimi K3 is genuine innovation performing near GPT-5.4–5.5 on coding; Thompson's claim that distillation scales with RL is factually incorrect; a US ban would harm American cybersecurity defenders who rely on Chinese open models for tasks closed US models refuse to perform.

Evolution: April predictions emphasized RL-era structural limits for open-weight models; July analysis concedes K3 exceeded expectations, adds a direct rebuttal of Thompson, and introduces a concrete cybersecurity defender harm as a policy argument.

Zvi Mowshowitz

Kimi K3 is real but four to six months behind the closed frontier; benchmark scores are systematically overstated due to max-effort evaluation, jagged capabilities, and Claude distillation; open-weight release carries roughly a 10% chance of consequential regret.

Evolution: Consistent skeptical-but-rigorous position throughout.

Jack Clark (Import AI) / UK AISI

The shrinking open-to-closed gap on cybersecurity tasks — now 4–7 months — means defenders have a short window before frontier cyber capabilities are broadly accessible; governance bodies deserve serious consideration.

Evolution: Consistent; UK AISI provided the first government-body quantification on a specific high-risk domain.

Ben Thompson (Stratechery)

The US should go on offense through legislation legalizing AI training data collection and barring distillation restrictions for US companies, rather than trying to restrict Chinese models.

Evolution: Consistent; Lambert's July 22 rebuttal of the distillation-scales-with-RL claim remains unanswered by Thompson.

Simon Willison

Kimi K3 marks the moment open-weight models can stand toe-to-toe with proprietary frontier ones; endorses Thompson's fair-use legislation; frames DeepSeek V4-Flash-0731 as further evidence and an Anthropic security incident as relevant policy context.

Evolution: Expanded coverage from Moonshot to DeepSeek and the Anthropic incident; consistent endorsement of open-weight development.

Trump administration / US Congress

Pursuing restrictions on Chinese AI broadly; reviving an AI model ban; Congress introduced kill-switch legislation; both efforts face severe enforcement limits against open weights; the Bloomberg compute report adds a potential direct enforcement case.

Evolution: Moved from theoretical policy debate to active government action; the Bloomberg H200 report adds a potential concrete enforcement avenue distinct from the open-weight distribution problem.

Alibaba / Moonshot AI

Both labs release frontier-scale models for broad developer adoption; Qwen3.8-Max remains API-only despite open-weight commitment; Alibaba denied supplying H200 chips to Moonshot but did not deny the supply arrangement and would not identify the chips.

Evolution: Alibaba reversed earlier reluctance to commit to open weights following Xi's endorsement but has not delivered; the Bloomberg compute report put Alibaba in a defensive posture handled through partial denial.

Florian Brand (Interconnects)

Predicted open-weight consolidation has failed to materialize; commercial viability of open-weight AI is now demonstrated; K3's commercial license restrictions could create regulatory leverage for US government action against American businesses using Chinese AI models.

Evolution: New voice this pass; adds an ecosystem-dynamics perspective and a licensing-leverage policy argument not previously in the thread.

Tensions

  • Lambert estimates the open-to-closed gap at 3–5 months; Mowshowitz puts it at 4–6 months; UK AISI measures 4–7 months for cybersecurity specifically — the floor matters for decisions about intervention timing. [2][3][4]
  • Semafor reports K3 underperforms top frontier models on cybersecurity vulnerability detection; Lambert rates K3 near GPT-5.4–5.5 on coding — suggesting a jagged capability profile or divergent evaluation methodology. [11][10]
  • Thompson argues distillation becomes more impactful as RL scales; Lambert explicitly calls this factually incorrect, arguing RL requires millions of rollouts prohibitively expensive to run through closed APIs; Brockman frames distillation as a technical problem with technical solutions. [19][10][11]
  • Lambert argues a ban on Chinese AI models would harm US cybersecurity defenders who rely on them for analysis closed models refuse to perform; the Trump administration is pursuing a ban on national security grounds. [10][14]
  • Thompson and Willison argue the US should go on offense through fair-use legislation; Hassabis argues for a pre-release testing body; Semafor argues the correct posture is simply to prepare for a world where powerful models are freely downloadable without guardrails. [19][1][4][11]
  • Bloomberg and multiple sources identify Moonshot's compute as H200s supplied by Alibaba; Alibaba denies the H200 identification but does not deny the supply arrangement and refuses to disclose what chips were provided — the core factual dispute is unresolved. [13]

Sources

  1. [1] moonshotai/Kimi-K3 — Simon Willison (2026-07-27)
  2. [2] Kimi K3: The open-weights escalation — Interconnects (2026-07-20)
  3. [3] On Kimi K3: Its Capabilities And Related Discontents — Zvi's AI Roundups (2026-07-20)
  4. [4] Import AI 465: Open vs closed gaps; Kimi K3; Demis' big policy plan — Import AI (2026-07-20)
  5. [5] 😸 Alibaba’s 2.4T Qwen joins the AI race — The Neuron (2026-07-20)
  6. [6] qwen3.8-max-preview - API Pricing & Performance | AIHubMix — reactive:kimi-k3-chinese-open-weights
  7. [7] SemiAnalysis tested Qwen3.8-Max-Preview 2.4T params https://t.co/iN84UKxB4b — SemiAnalysis Twitter (2026-08-02)
  8. [8] Qwen 3.8 Max just built better 3D physics scenes than Fable 5 while costing about 7x less to run. — Rohan Paul Twitter (2026-08-03)
  9. [9] deepseek-ai/DeepSeek-V4-Flash-0731 — Simon Willison (2026-07-31)
  10. [10] Open models recap: more on Kimi K3, Qwen 3.8, Xi's WAIC speech, distillation, the open-closed gap, and what's next — Interconnects (2026-07-22)
  11. [11] 🟡 Kill switch — Semafor Technology (2026-07-24)
  12. [12] Latest open artifacts (#23): Laguna S2.1, Inkling, & Kimi K3 show the utility of open models on the Pareto frontier — Interconnects (2026-08-02)
  13. [13] Bloomberg reports that Moonshot runs its Kimi models on roughly 20,000 Nvidia chips it accesses through a confidential d… — Rohan Paul Twitter (2026-07-31)
  14. [14] Trump administration reportedly reviving push to ban Chinese AI models, Kimi K3 — reactive:kimi-k3-chinese-open-weights (2026-07-21)
  15. [15] 😺 3,000 Mexican exam scores wiped over AI — The Neuron (2026-08-03)
  16. [16] Really cool. @atomic_chat_hq 's 1-bit Kimi K3 quant shrinks the 2.8T model to 590GB (-62%) — Rohan Paul Twitter (2026-07-31)
  17. [17] Oxide and Friends: The Open Weight Revolution with Simon Willison — Simon Willison (2026-07-31)
  18. [18] My bets on open models, mid-2026 — Interconnects (2026-04-15)
  19. [19] Who’s Afraid of Chinese Models? — Simon Willison (2026-07-20)