Google DeepMind DiffusionGemma: Parallel Diffusion Architecture for 4x Faster Local Text Generation
Synthesis history
4 versions, newest first.
-
Version 4 2026-06-16 18:24 UTC · 62 items
The substantive new development is DiffusionGemma's appearance in Google Cloud's Agent Platform Model Garden [^30193], indicating integration into Google's own cloud infrastructure beyond the initial Hugging Face and NV…
-
Version 3 2026-06-14 02:30 UTC · 57 items
The one substantive new development is a third-party benchmark from atomic[.]chat on local H100 hardware (FP8) confirming the 4x speed claim against Gemma 4 26B A4B [^28369], adding an independent non-cloud data point t…
-
Version 2 2026-06-12 02:06 UTC · 28 items
Two additional voices joined the coverage: Ars Technica confirmed core architectural claims and framed the model as practical for local deployment [^27904], and Simon Willison identified DiffusionGemma as a public retur…
-
Version 1 2026-06-10 18:10 UTC · 15 items
Google DeepMind released DiffusionGemma on June 10, 2026, a 26B mixture-of-experts text model that generates up to 256 tokens simultaneously by denoising from noise rather than predicting one token at a time. [^27862] T…