The Information Machine

Kimi K3 and Qwen3.8-Max: Chinese Labs Close Gap with Closed Frontier Models · history

Version 3

2026-07-23 08:05 UTC · 68 items

What

Moonshot AI's Kimi K3 (2.8 trillion parameters, open weights) and Alibaba's Qwen3.8-Max (2.4 trillion parameters, open-weight release committed) have brought Chinese models to within 3–7 months of closed frontier systems by most analyst estimates [2][5][6]. Nathan Lambert's July 22 follow-up places Kimi K3 at roughly GPT-5.4–5.5 on coding tasks and argues Chinese gains reflect genuine innovation rather than distillation-driven copying, directly challenging Ben Thompson's framing [1]. The Trump administration is reportedly pursuing a ban on Chinese AI models, but the debate has expanded: Lambert argues such a ban would hurt US cybersecurity defenders who rely on these models for analysis tasks closed US models refuse to perform [8][1].

Why it matters

Chinese labs are now releasing near-frontier open-weight models on 1–2 month cycles [1], compressing the time before high-capability AI is broadly accessible. The US policy options are doubly constrained: open-weight distribution makes a ban structurally difficult to enforce, and restricting these models may disadvantage American defenders relative to attackers globally who face no such restriction.

Open questions

  • What enforcement mechanism could a Trump administration Chinese AI ban use that would be effective against open weights already available for download worldwide? [8]

  • When will Alibaba release Qwen3.8-Max as open weights, and will independent benchmarks confirm its self-reported claim of trailing only Claude Fable 5? [3]

  • Lambert argues distillation has limited RL impact because running millions of rollouts through closed-model APIs is prohibitively expensive — does empirical research confirm or refute this claim, and does it undercut Thompson's policy argument? [1]

  • If a ban on Chinese AI models is enacted, does harm to US cybersecurity defenders who rely on them for analysis closed US models won't perform outweigh the security benefits? [1]

Narrative

Moonshot AI's Kimi K3, a 2.8 trillion-parameter mixture-of-experts model released as open weights in July 2026, is assessed by multiple analysts as the strongest open-weight model to date — performing at approximately GPT-5.4 to 5.5 level on coding tasks with strong agentic research capabilities [1][2]. Alibaba's Qwen3.8-Max (2.4 trillion parameters) was previewed with self-reported benchmarks claiming it trails only Claude Fable 5, with open-weight release committed following Xi Jinping's July 20 WAIC keynote publicly endorsing open-source AI [3][4]. Multiple analysts place the open-to-closed capability gap at 3–7 months, down from 6–10 months in prior periods [2][5][6]. Independent testing placed Kimi K3 fourth among all evaluated models [7].

A substantive debate has formed around distillation's role in Chinese model gains. Nathan Lambert argues Chinese labs' performance is genuine innovation: RL training — where he contends real capability gains occur — requires millions of rollouts that would be prohibitively expensive and slow to run through closed-model APIs, meaning distillation from those APIs has limited practical impact on model quality [1]. He characterizes these as models people compare favorably on internal training benchmarks, not distillation artifacts, and explicitly calls Ben Thompson's argument — that distillation becomes more impactful as RL scales — factually incorrect and potentially misleading to policymakers [1]. Zvi Mowshowitz holds a different position: Kimi K3 benchmarks are systematically inflated by max-effort evaluation settings, jagged capability profiles, and some Claude distillation, though he concedes the model is real and roughly four to six months behind the closed frontier [6].

The Trump administration is reportedly reviving a push to ban Chinese AI models following Kimi K3's launch, citing cybersecurity concerns, with officials acknowledging that open-weight distribution makes a comprehensive ban nearly unenforceable [8]. Lambert introduces a concrete counterargument: Hugging Face was able to analyze the OpenAI breach only by using a Chinese open model, because US frontier models had guardrails blocking the analysis — meaning a ban would directly disadvantage American cybersecurity defenders relative to global attackers who face no equivalent restriction [1]. This extends the enforcement problem from structural (open weights cannot be recalled) to strategic (restricting these tools harms US security). Ben Thompson's earlier proposal — legislation making AI training data collection explicit fair use and barring distillation restrictions for US companies — represents an offensive alternative posture, while Demis Hassabis proposes a FINRA-style pre-release standards body [4][5].

Chinese labs' cost efficiency advantage stems from focused engineering teams with no distracting side projects, lower operational costs, and increasing access to domestic compute including Huawei Ascend chips [1]. The open-weight ecosystem has professionalized to the point where Chinese labs now release models on 1–2 month cycles matching closed lab iteration speeds [1]. Nathan Lambert's April 2026 structural prediction — that RL-era agentic deployment data would be the first domain where closed labs hold a durable edge, and that Chinese open-weight labs would face funding difficulties by late 2026 [9] — remains unresolved: K3 and Qwen3.8-Max are the first major tests of that hypothesis, and Lambert's own July 22 analysis concedes K3 exceeded expectations on closing the capability gap.

Timeline

  • 2026-04-15: Nathan Lambert publishes predictions arguing RL-era agentic deployment data is the first domain where closed labs can durably outperform open-weight models, and that Chinese open-weight labs will face funding difficulties by late 2026. [9]
  • 2026-07-20: Xi Jinping delivers WAIC keynote publicly endorsing open-source AI, prompting Alibaba to commit to releasing Qwen3.8-Max as open weights after earlier hesitation. [4][2]
  • 2026-07-20: Moonshot AI releases Kimi K3 as open weights — a 2.8 trillion-parameter MoE model assessed by multiple analysts as the strongest open-weight model yet released, approaching but below Claude Fable 5 and GPT 5.6 Sol. [5][2][6]
  • 2026-07-20: Alibaba previews Qwen3.8-Max (2.4 trillion parameters) with self-reported benchmarks claiming it trails only Claude Fable 5; open-weight release committed but not yet delivered. [3]
  • 2026-07-20: UK AI Security Institute reports the capability gap between open-weight and closed frontier models on cybersecurity tasks has narrowed from 6–10 months to 4–7 months. [5]
  • 2026-07-20: Ben Thompson proposes US legislation making AI training data collection explicit fair use and barring distillation restrictions for US companies as an offensive alternative to restricting open-weight models. [4]
  • 2026-07-20: Demis Hassabis proposes a FINRA-style public-private standards body to test frontier AI systems before release, with eventual mandatory compliance. [5]
  • 2026-07-21: Trump administration is reportedly reviving efforts to ban Chinese AI models following Kimi K3's launch; officials note open weights make an outright ban nearly impossible to enforce. [8]
  • 2026-07-21: Kimi K3 places fourth in independent AI model testing across evaluated models. [7]
  • 2026-07-22: Nathan Lambert publishes follow-up arguing distillation has limited RL impact, explicitly calling Thompson's distillation narrative factually incorrect, and introducing the cybersecurity defender argument against a ban. [1]

Perspectives

Nathan Lambert (Interconnects)

Kimi K3 is genuine innovation — not distillation — performing near GPT-5.4–5.5 on coding; Thompson's claim that distillation scales with RL is factually incorrect; a US ban would harm American cybersecurity defenders more than it restricts attackers.

Evolution: April predictions emphasized RL-era structural limits and funding difficulties; July 22 follow-up concedes K3 exceeded expectations, adds a direct rebuttal of Thompson, and introduces the cybersecurity defender harm as a concrete policy argument.

Zvi Mowshowitz

Kimi K3 is real but four to six months behind the closed frontier; benchmark scores are systematically overstated due to max-effort evaluation, jagged capabilities, and Claude distillation; open-weight release carries roughly a 10% chance of consequential regret.

Evolution: Consistent skeptical-but-rigorous position that pushes back on both overclaims and underclaims.

Jack Clark (Import AI)

The shrinking open-to-closed gap on cybersecurity tasks — now 4–7 months per UK AISI — means cyber defenders have a short window before frontier cyber capabilities become broadly accessible; governance bodies deserve serious consideration.

Evolution: Consistent concern about open-weight proliferation; first appearance with specific cybersecurity domain quantification from a government body.

Ben Thompson (Stratechery) / Simon Willison

Thompson proposes US legislation legalizing training data collection and barring distillation restrictions as an offensive alternative to banning Chinese models; Willison endorses this as pragmatic and frames existing US lab policies as hypocritical.

Evolution: Consistent; Lambert's July 22 piece now directly challenges Thompson's distillation-scales-with-RL claim as factually incorrect.

Trump administration

Reportedly reviving a push to ban Chinese AI models following Kimi K3, citing cybersecurity concerns; administration officials acknowledge open-weight distribution makes a comprehensive ban nearly unenforceable.

Evolution: Moved the policy debate from theoretical proposals to an active government consideration.

Alibaba / Moonshot AI (Chinese labs)

Both labs are releasing frontier-scale models as open weights on 1–2 month cycles, pursuing dual goals of maximizing raw capability while enabling broad developer adoption.

Evolution: Alibaba reversed earlier reluctance to release Qwen3.8-Max as open weights following Xi Jinping's public endorsement.

UK AI Security Institute

Empirically documents that the open-to-closed capability gap on cybersecurity tasks has narrowed to 4–7 months, down from a prior range of 6–10 months.

Evolution: First quantitative government-body measurement on a specific high-risk domain in this thread.

Ethan Mollick

The gap between open and closed models is larger than benchmarks suggest; open models are more fragile on out-of-distribution problems and exhibit lower emergent capabilities than standard evaluations capture.

Evolution: Counter-voice to open-weight optimism; focuses on real-world fragility rather than benchmark scores.

Tensions

  • Lambert estimates the open-to-closed capability gap at 3–5 months; Mowshowitz puts it at 4–6 months; UK AISI measures 4–7 months for cybersecurity specifically — the floor matters for decisions about intervention timing. [2][6][5]
  • Mowshowitz argues Kimi K3 benchmarks are systematically inflated by max-effort evaluation, jagged capability profiles, and Claude distillation; Lambert argues these are genuinely good models validated on internal training benchmarks, not distillation artifacts. [6][2][1]
  • Thompson argues distillation becomes more impactful as RL scales; Lambert explicitly calls this factually incorrect, arguing RL training requires millions of rollouts prohibitively expensive to run through closed APIs. [4][1]
  • The Trump administration is reportedly pursuing a ban on Chinese AI models, but Lambert argues such a ban would harm US cybersecurity defenders who rely on these models for analysis tasks closed US models refuse to perform. [8][1]
  • Thompson and Willison argue the US should go on offense through fair-use legislation, while Hassabis argues for a pre-release testing body — reflecting different assumptions about whether openness or governance is the better strategic posture. [4][5]
  • Lambert argues RL-era agentic deployment data is the first durable structural edge closed labs hold over open-weight models; Mollick argues the gap is broadly larger than benchmarks show across multiple dimensions including out-of-distribution robustness. [9][10]

Sources

  1. [1] Open models recap: more on Kimi K3, Qwen 3.8, Xi's WAIC speech, distillation, the open-closed gap, and what's next — Interconnects (2026-07-22)
  2. [2] Kimi K3: The open-weights escalation — Interconnects (2026-07-20)
  3. [3] 😸 Alibaba’s 2.4T Qwen joins the AI race — The Neuron (2026-07-20)
  4. [4] Who’s Afraid of Chinese Models? — Simon Willison (2026-07-20)
  5. [5] Import AI 465: Open vs closed gaps; Kimi K3; Demis' big policy plan — Import AI (2026-07-20)
  6. [6] On Kimi K3: Its Capabilities And Related Discontents — Zvi's AI Roundups (2026-07-20)
  7. [7] Kimi K3 reaches fourth place in independent AI testing — reactive:kimi-k3-chinese-open-weights
  8. [8] Trump administration reportedly reviving push to ban Chinese AI models, Kimi K3 — reactive:kimi-k3-chinese-open-weights (2026-07-21)
  9. [9] My bets on open models, mid-2026 — Interconnects (2026-04-15)
  10. [10] Ethan Mollick on X: "This is a good explanation of why the gap between open and closed models is larger than it appears in benchmarks. I would add in that current open models are also more fragile than closed: they handle out-of-distribution problems far less well & have lower emergent capabilities." / X — reactive:kimi-k3-chinese-open-weights