Google DeepMind Launches Gemini Robotics 2 and ER 2 in Wave of Physical AI Releases · history
Version 2
2026-08-01 02:08 UTC · 32 items
What
Google DeepMind released Gemini Robotics 2.0 in late July 2026, comprising three sub-models: Gemini Robotics ER 2 (publicly available via the Gemini Live API), a vision-language-action model, and an on-device model for hardware-constrained deployment [2][1]. The system provides whole-body humanoid control, on-device adaptation from under 200 examples, and multi-robot collaboration across heterogeneous platforms [2]. Separately, Black Forest Labs launched FLUX 3 around July 25 — a single foundation model trained jointly on images, video, audio, and robotic actions — with robotics demonstrated on Audi factory tasks [7]. DeepMind also released the ASIMOV-Agentic benchmark, backed by peer-reviewed papers and an official site, to test whether foundation model orchestrators can refuse unsafe robot commands [4][5][6].
Why it matters
Physical AI is moving from demos toward deployable systems, with DeepMind claiming production-grade latency via a public API and on-device operation without network connectivity. Whether the ASIMOV-Agentic benchmark gains adoption beyond DeepMind's own evaluations will determine whether there is a shared standard for robot safety — or whether each lab defines safety on its own terms.
Open questions
When will Gemini Robotics VLA and On-Device 2 move beyond early-access partners to general availability? [2]
Will the ASIMOV-Agentic safety benchmark gain adoption beyond DeepMind's own models as an industry evaluation standard? [6][4]
Can FLUX 3's claim that generative video training builds physical world understanding hold up beyond self-reported comparisons against Runway and Luma? [7]
How do DeepMind's specialized layered models and FLUX 3's unified approach compare on independent, third-party manipulation benchmarks?
Narrative
Google DeepMind released Gemini Robotics 2.0 in late July 2026, describing its goal as building a 'generalist robot' capable of executing arbitrary human-directed tasks — what its scientists call 'physical AGI' [1]. The release covers three interconnected models. Gemini Robotics 2 provides whole-body humanoid control from feet to fingertips, enabling robots to walk, crouch, and manipulate objects in unstructured environments [2][1]. Gemini Robotics On-Device 2 is designed for hardware-constrained deployment: it adapts to new robot morphologies in a few hours using fewer than 200 demonstration examples, running locally without network access [2]. Gemini Robotics ER 2, the embodied reasoning layer, achieves 91.3% accuracy on moment-finding tasks at 4x the speed of competing larger models, and integrates with the Gemini Live API for sub-second latency in real-time task orchestration [3]. Of the three, only ER 2 is publicly available; the VLA and On-Device models remain limited to early-access partners [2].
A notable feature across the stack is multi-robot collaboration: heterogeneous robots — humanoids alongside bi-arm platforms — communicate via shared semantic understanding to hand off subtasks no single robot could complete alone [3][2]. On safety, DeepMind introduced the ASIMOV-Agentic benchmark, which tests whether high-level foundation model orchestrators can refuse unsafe tool calls from lower-level action models and seek human clarification when physical feasibility is uncertain [3][2]. The benchmark is backed by peer-reviewed work in MLRS proceedings and on arXiv, with an official site at asimov-benchmark.github.io [4][5][6].
Black Forest Labs took a different architectural approach with FLUX 3, launched around July 25. Rather than a specialized robotics stack, FLUX 3 is a single model trained jointly on images, video, audio, and robotic actions, with robotics treated as a downstream application of generative media training rather than a separate engineering effort [7]. The model was demonstrated on Audi production-line tasks involving cables, seals, and deformable parts that rule-based automation handles poorly [7]. Black Forest Labs reported that FLUX 3 outperformed Runway Gen-4.5 in 77% of comparisons and Luma Ray 3.2 in 93%, though these figures come from internal evaluations [7]. The underlying argument is that generating convincing video of physical tasks requires internalizing causal rules about how the physical world behaves, making generative video and robot action training mutually reinforcing.
Timeline
- 2026-07-25: Black Forest Labs launches FLUX 3, a unified model spanning images, video, audio, and robotic actions, with an Audi factory robotics demo. [8][9][10]
- 2026-07-27: The Neuron covers FLUX 3, framing the convergence of generative video and robotics training as a shared architectural foundation for physical intelligence. [7]
- 2026-07-28: Google DeepMind publishes Gemini Robotics 2, covering full-body humanoid control, on-device adaptation under 200 examples, and multi-robot collaboration. [2]
- 2026-07-30: Google DeepMind publishes Gemini Robotics ER 2 and makes it available to developers via the Gemini Live API; Ars Technica reports on the 'physical AGI' framing from DeepMind scientists. [3][1]
Perspectives
Google DeepMind
Presents Gemini Robotics 2.0 as a step toward 'physical AGI' — a generalist robot capable of arbitrary human-directed tasks — with a layered stack (reasoning, VLA, on-device), a peer-reviewed safety benchmark in ASIMOV-Agentic, and tiered developer access starting with ER 2.
Evolution: Consistent with DeepMind's prior Gemini Robotics line; this release adds whole-body control, multi-robot heterogeneous collaboration, and an explicit safety evaluation framework.
Black Forest Labs
Argues for a unified multimodal foundation model covering image, video, audio, and action in a single architecture, with robotics as a downstream application of generative media training rather than a separate stack.
Evolution: New entrant to physical AI; FLUX 3 extends the company's prior image generation work into video and robotics for the first time.
The Neuron (Grant Harvey)
Bullish on the convergence thesis: training models to generate convincing physical-world video and training robots to act in the physical world share underlying causal requirements, making generative video a viable route to physical intelligence.
Evolution: Consistent newsletter framing applied to both FLUX 3 and the broader wave of physical AI releases.
Ars Technica (Ryan Whitwam)
Reports DeepMind's capability claims straightforwardly, relaying the 'physical AGI' framing without significant skepticism or independent technical evaluation.
Evolution: First appearance in this thread; provides mainstream third-party corroboration of the announcement.
Tensions
- DeepMind builds specialized, layered robotics models (reasoning, VLA, on-device as separate tiers); Black Forest Labs argues a single unified multimodal model trained on video and action is sufficient and more efficient. [3][2][7]
- DeepMind makes ER 2 publicly available but keeps VLA and On-Device 2 behind early-access restrictions, limiting independent verification of its most capable physical control models. [2][1]
- FLUX 3's performance claims rest on self-reported internal evaluations against Runway and Luma, while DeepMind publishes structured benchmark results backed by peer-reviewed papers, making direct cross-system comparison unavailable. [7][3][4][5]
Sources
- [1] Google reveals Gemini Robotics 2.0, promising improved dexterity and safety — Ars Technica AI (2026-07-30)
- [2] Gemini Robotics 2 brings whole body intelligence to robots — DeepMind Blog (2026-07-28)
- [3] Gemini Robotics ER 2: powering robotics with video understanding, task orchestration, and multi-robot collaboration — DeepMind Blog (2026-07-30)
- [4] Generating Robot Constitutions & Benchmarks for Semantic Safety — reactive:gemini-robotics-2-launch
- [5] Generating Robot Constitutions & Benchmarks for Semantic Safety — reactive:gemini-robotics-2-launch
- [6] ASIMOV Benchmark v1 — reactive:gemini-robotics-2-launch
- [7] 😸 Multimodal AI just got real — The Neuron (2026-07-27)
- [8] FLUX 3 Launches: Black Forest Labs Enters Video, Audio, and Physical AI in One Model — reactive:gemini-robotics-2-launch
- [9] FLUX 3 x mimic: The Next Generation of Video-Action Models | Black Forest Labs — reactive:gemini-robotics-2-launch
- [10] Black Forest Labs Unveils First Model for Robotics in Shift to Physical AI — reactive:gemini-robotics-2-launch