Kimi K3 and Qwen3.8-Max: Chinese Labs Close Gap with Closed Frontier Models · history
Version 2
2026-07-22 02:10 UTC · 40 items
What
Moonshot AI's Kimi K3 (2.8 trillion parameters, open weights released) and Alibaba's Qwen3.8-Max (2.4 trillion parameters, open-weight release committed) have brought Chinese models to within 3–7 months of closed frontier systems by most analyst estimates [1][4][5]. Independent testing placed Kimi K3 fourth among all models evaluated [6]. The Trump administration is reportedly reviving a push to ban Chinese AI models following Kimi K3's launch, citing cybersecurity concerns, though officials acknowledge that open-weight distribution makes a comprehensive ban nearly impossible to enforce [8].
Why it matters
Open-weight models at near-frontier capability compress the time before high-capability AI is broadly available — including for adversarial use — while the open-weights distribution mechanism makes after-the-fact government restrictions structurally difficult to enforce. The Trump administration's reported ban consideration, combined with analysts' finding that restricting domestic access does nothing to limit global availability of Chinese open-weight alternatives, defines the core policy bind.
Open questions
What enforcement mechanism could a Trump administration Chinese AI ban use that would be effective against open weights already available for download worldwide? [8]
When will Alibaba release Qwen3.8-Max as open weights, and will independent benchmarks confirm its self-reported claim of trailing only Claude Fable 5? [3]
Do independent benchmark results for Kimi K3 support or contradict Mowshowitz's claim that evaluations are systematically optimistic due to max-effort settings and jagged capability profiles? [5][6][7]
Does the online reinforcement-learning advantage — where closed labs hold exclusive access to real-world agentic deployment data — create a durable capability gap open-weight models cannot close, as Lambert predicted in April? [9][1]
Narrative
Moonshot AI's Kimi K3, a 2.8 trillion-parameter mixture-of-experts model, is released as open weights and is assessed by multiple analysts as the strongest open-weight model released to date [1]. Alibaba's Qwen3.8-Max, at 2.4 trillion parameters, is previewed and committed to open-weight release following Xi Jinping's WAIC keynote publicly endorsing open-source AI — a signal both Simon Willison and Nathan Lambert read as the Chinese government's strategic assessment that current frontier models do not pose unacceptable national-security risks [2][1][3]. Nathan Lambert, Jack Clark, and Zvi Mowshowitz broadly agree the gap between these models and leading closed systems — Claude Fable 5 and GPT 5.6 Sol — is roughly three to seven months, down from six to ten months in prior periods [1][4][5]. The UK AI Security Institute quantifies the gap on cybersecurity tasks specifically at four to seven months [4]. Independent testing placed Kimi K3 fourth among all models evaluated [6].
There is genuine disagreement about how to read Kimi K3's results. Mowshowitz argues benchmark scores are systematically optimistic: evaluations were run at maximum effort, the model has jagged and uneven capabilities, and it has been substantially distilled from Claude models — accounting for some but not all measured gains [5]. He also notes K3 sits on the established Chinese capability trend line rather than above it, and that recurring narratives of Chinese AI erasing America's lead follow a pattern visible since the original DeepSeek moment [5]. Commentary on "the catch behind Kimi K3's benchmark leap" suggests the skeptical reading has traction beyond Mowshowitz alone [7]. Lambert reads the same results differently: Chinese labs are achieving near-frontier performance with orders of magnitude less funding than US counterparts, and high-capability open-weight releases are economically decelerationist for frontier labs — compressing margins and reducing terminal valuations while accelerating broader AI diffusion [1].
The policy debate has moved from theoretical to active. The Trump administration is reportedly reviving efforts to ban Chinese AI models following Kimi K3's launch, citing cybersecurity concerns, but officials and analysts acknowledge that open-weight distribution makes an outright US ban nearly impossible to enforce once weights are available for download [8]. This enforcement problem underscores the asymmetry Lambert and Clark had already identified: domestic restrictions would disadvantage US developers without meaningfully reducing global access to Chinese open-weight alternatives [1][4]. Ben Thompson's counterproposal — legislation making AI training data collection explicit fair use and barring distillation restrictions for US companies — represents a different posture, going on offense rather than defense [2]. Demis Hassabis separately proposes a FINRA-style public-private standards body to evaluate frontier AI systems before release [4].
Nathan Lambert's April 2026 analysis provides longer-range context: he predicted that the reinforcement-learning-dominated training era would create the first durable domain where closed labs can outperform open-weight models — specifically, real-world agentic deployment data for online RL — and that Chinese open-weight labs would begin facing funding difficulties as soon as late 2026, with capability trajectory divergence visible three to nine months later [9]. Whether K3 and Qwen3.8-Max confirm or complicate that prediction will shape how analysts read the open-weight trajectory through the rest of the year.
Timeline
- 2026-04-15: Nathan Lambert publishes structured predictions arguing that RL-era agentic deployment data is the first domain where closed labs can durably outperform open-weight models, and that Chinese open-weight labs will face funding difficulties by late 2026. [9]
- 2026-07-20: Xi Jinping delivers WAIC keynote publicly endorsing open-source AI, prompting Alibaba to commit to releasing Qwen3.8-Max as open weights after earlier hesitation. [2][1]
- 2026-07-20: Moonshot AI releases Kimi K3 as open weights — a 2.8 trillion-parameter MoE model assessed by multiple analysts as the strongest open-weight model yet released, approaching but below Claude Fable 5 and GPT 5.6 Sol. [4][1][5]
- 2026-07-20: Alibaba previews Qwen3.8-Max (2.4 trillion parameters) with self-reported benchmarks claiming it trails only Claude Fable 5; open-weight release committed but not yet delivered. [3]
- 2026-07-20: UK AI Security Institute reports the capability gap between open-weight and closed frontier models on cybersecurity tasks has narrowed from 6–10 months to 4–7 months. [4]
- 2026-07-20: Ben Thompson proposes US legislation making AI training data collection explicit fair use and barring distillation restrictions for US companies, as an offensive alternative to restricting open-weight models. [2]
- 2026-07-20: Demis Hassabis proposes a FINRA-style public-private standards body to test frontier AI systems before release, with eventual mandatory compliance. [4]
- 2026-07-21: Trump administration is reportedly reviving efforts to ban Chinese AI models following Kimi K3's launch, citing cybersecurity concerns; officials note open weights make an outright ban nearly impossible to enforce. [8]
- 2026-07-21: Kimi K3 places fourth in independent AI testing across evaluated models. [6]
Perspectives
Nathan Lambert (Interconnects)
Kimi K3 reduces the open-to-closed gap to roughly 3–5 months; Chinese capital efficiency is the underlying structural story; US restrictions would create asymmetric disadvantage without meaningfully reducing risk.
Evolution: April predictions emphasized RL-era structural limits and funding difficulties as binding constraints; July analysis confirms those frames but concedes K3 exceeded expectations on closing the capability gap.
Zvi Mowshowitz
Kimi K3 is real but four to six months behind the closed frontier; benchmark scores are systematically overstated due to max-effort evaluation, jagged capabilities, and Claude distillation; open-weight release carries roughly a 10% chance of consequential regret.
Evolution: Consistent skeptical-but-rigorous position that pushes back on both overclaims and underclaims.
Jack Clark (Import AI)
The shrinking open-to-closed gap on cybersecurity tasks — now 4–7 months per UK AISI — means cyber defenders have a short window before frontier cyber capabilities become broadly accessible; governance bodies deserve serious consideration.
Evolution: Consistent concern about open-weight proliferation; first appearance with specific cybersecurity domain quantification from a government body.
Simon Willison
US labs are hypocritical in forbidding distillation via terms of service while training on unlicensed data; Ben Thompson's proposal to legalize training data collection and bar distillation restrictions for US companies is a pragmatic alternative to restriction.
Evolution: Consistent; frames the issue as a policy contradiction rather than a capability or safety debate.
Trump administration
Reportedly reviving a push to ban Chinese AI models following Kimi K3, citing cybersecurity concerns; administration officials acknowledge open-weight distribution makes a comprehensive ban nearly impossible to enforce.
Evolution: First appearance on this thread; moves the policy debate from theoretical proposals to an active government consideration.
Alibaba / Moonshot AI (Chinese labs)
Both labs are releasing frontier-scale models as open weights, pursuing dual goals: maximizing raw capability while making it cheap and accessible for broad developer adoption.
Evolution: Alibaba reversed earlier reluctance to release Qwen3.8-Max as open weights following Xi Jinping's public endorsement of open-source AI.
UK AI Security Institute
Empirically documents that the open-to-closed capability gap on cybersecurity tasks has narrowed to 4–7 months, down from a prior range of 6–10 months.
Evolution: First quantitative government-body measurement on a specific high-risk domain in this thread.
Ethan Mollick
The gap between open and closed models is larger than benchmarks suggest; open models are more fragile on out-of-distribution problems and exhibit lower emergent capabilities than standard evaluations capture.
Evolution: Counter-voice to open-weight optimism; focuses on real-world fragility rather than benchmark scores.
Tensions
- Lambert estimates the open-to-closed capability gap at 3–5 months; Mowshowitz puts it at 4–6 months; UK AISI measures 4–7 months for cybersecurity specifically — the floor matters for decisions about intervention timing. [1][5][4]
- Mowshowitz argues Kimi K3 benchmarks are systematically inflated by max-effort evaluation, jagged capability profiles, and Claude distillation; Lambert and Clark treat the benchmarks as meaningful evidence of genuine capability gains. [5][1][4]
- The Trump administration is reportedly pursuing a ban on Chinese AI models, but officials acknowledge open-weight distribution makes such a ban nearly unenforceable — a structural contradiction at the center of the current policy situation. [8][1][4]
- Thompson and Willison argue the US should go on offense — legalizing training data collection and barring distillation restrictions — while Hassabis argues for a pre-release testing body, reflecting different assumptions about whether openness or governance is the better strategic posture. [2][4]
- Lambert argues RL-era agentic deployment data is the first durable structural edge closed labs hold over open-weight models; Mollick argues the gap is broadly larger than benchmarks show across multiple dimensions including out-of-distribution robustness. [9][10]
Sources
- [1] Kimi K3: The open-weights escalation — Interconnects (2026-07-20)
- [2] Who’s Afraid of Chinese Models? — Simon Willison (2026-07-20)
- [3] 😸 Alibaba’s 2.4T Qwen joins the AI race — The Neuron (2026-07-20)
- [4] Import AI 465: Open vs closed gaps; Kimi K3; Demis' big policy plan — Import AI (2026-07-20)
- [5] On Kimi K3: Its Capabilities And Related Discontents — Zvi's AI Roundups (2026-07-20)
- [6] Kimi K3 reaches fourth place in independent AI testing — reactive:kimi-k3-chinese-open-weights
- [7] The catch behind Kimi K3's benchmark leap | The Deep View — reactive:kimi-k3-chinese-open-weights
- [8] Trump administration reportedly reviving push to ban Chinese AI models, Kimi K3 — reactive:kimi-k3-chinese-open-weights (2026-07-21)
- [9] My bets on open models, mid-2026 — Interconnects (2026-04-15)
- [10] Ethan Mollick on X: "This is a good explanation of why the gap between open and closed models is larger than it appears in benchmarks. I would add in that current open models are also more fragile than closed: they handle out-of-distribution problems far less well & have lower emergent capabilities." / X — reactive:kimi-k3-chinese-open-weights