Google DeepMind Launches Gemini Robotics 2 and ER 2 in Wave of Physical AI Releases · history
Version 3
2026-08-03 09:24 UTC · 47 items
What
Google DeepMind released Gemini Robotics 2.0 in late July 2026, comprising three sub-models: Gemini Robotics ER 2 (publicly available via the Gemini Live API), a vision-language-action model, and an on-device model for hardware-constrained deployment [2][1]. The release covers whole-body humanoid control, on-device adaptation from under 200 examples, and multi-robot collaboration across heterogeneous platforms [2]. DeepMind also released an accompanying safety tech report and the ASIMOV-Agentic benchmark, now at version 2, designed to test whether orchestrating models can refuse unsafe robot commands [7][4]. Black Forest Labs took a different approach with FLUX 3, a single model trained jointly on images, video, audio, and robotic actions, with robotics as a downstream application of generative media training rather than a separate stack [9].
Why it matters
Physical AI is moving from demos toward deployable systems, but no fully independent third-party benchmark has yet verified zero-shot transfer of Gemini Robotics 2 to arbitrary platforms [10], leaving capability claims primarily supported by DeepMind's own evaluations. Whether ASIMOV-Agentic becomes an industry-shared safety standard or remains a self-referential benchmark will determine whether robot safety has common ground across labs.
Open questions
When will Gemini Robotics VLA and On-Device 2 move beyond early-access partners to general availability? [2]
Will ASIMOV-Agentic gain adoption beyond DeepMind's own models as a shared industry safety evaluation standard? [4][8]
No fully independent third-party benchmark has shown reliable zero-shot transfer of Gemini Robotics 2 to arbitrary platforms — when will such verification emerge? [10]
Can FLUX 3's unified generative training approach match DeepMind's specialized robotics stack on independent manipulation benchmarks? [9]
Narrative
Google DeepMind released Gemini Robotics 2.0 in late July 2026, framing its goal as building a 'generalist robot' capable of executing arbitrary human-directed tasks — what its scientists call 'physical AGI' [1]. The release covers three interconnected models. Gemini Robotics 2 provides whole-body humanoid control from feet to fingertips, enabling robots to walk, crouch, and manipulate objects in unstructured environments [2][1]. Gemini Robotics On-Device 2 adapts to new robot morphologies in a few hours using fewer than 200 demonstration examples, running locally without network access [2]. Gemini Robotics ER 2, the embodied reasoning layer, achieves 91.3% accuracy on moment-finding tasks at 4x the speed of competing larger models, and integrates with the Gemini Live API for sub-second latency [3]. Of the three, only ER 2 is publicly available; the VLA and On-Device models remain limited to early-access partners [2].
DeepMind paired the release with a formal safety framework. The ASIMOV-Agentic benchmark, now at version 2, tests whether high-level foundation model orchestrators can refuse unsafe tool calls from lower-level action models and seek human clarification when physical feasibility is uncertain [4][5][6]. DeepMind published an official safety tech report alongside the release [7], and researcher Anirudha Majumdar has publicly promoted the benchmark as covering unsafe task refusal, uncertainty quantification, and human escalation [8]. The benchmark is backed by peer-reviewed work in MLRS proceedings and on arXiv [5][6]. A notable capability across the stack is multi-robot collaboration: heterogeneous robots communicate via shared semantic understanding to hand off subtasks no single robot could complete alone [3][2].
Black Forest Labs took a different architectural approach with FLUX 3, launched around July 25. Rather than a specialized robotics stack, FLUX 3 is a single model trained jointly on images, video, audio, and robotic actions, with robotics treated as a downstream application of generative media training [9]. The model was demonstrated on Audi production-line tasks involving cables, seals, and deformable parts that rule-based automation handles poorly [9]. Black Forest Labs reported FLUX 3 outperformed Runway Gen-4.5 in 77% of comparisons and Luma Ray 3.2 in 93%, though these figures come from internal evaluations [9].
One gap runs across both systems: no fully independent third-party benchmark has yet shown reliable zero-shot transfer of Gemini Robotics 2 to arbitrary platforms [10], and FLUX 3's performance comparisons rest on self-reported internal evaluations. DeepMind's structured benchmark results are backed by peer-reviewed papers, but the most capable physical control models — VLA and On-Device 2 — remain behind early-access restrictions, limiting external scrutiny [2][1]. The question of whether any lab's safety or capability claims will be independently stress-tested is unresolved.
Timeline
- 2026-07-25: Black Forest Labs launches FLUX 3, a unified model spanning images, video, audio, and robotic actions, with an Audi factory robotics demo. [11][12][13]
- 2026-07-27: The Neuron covers FLUX 3, framing the convergence of generative video and robotics training as a shared architectural foundation for physical intelligence. [9]
- 2026-07-28: Google DeepMind publishes Gemini Robotics 2, covering full-body humanoid control, on-device adaptation under 200 examples, and multi-robot collaboration. [2]
- 2026-07-30: Google DeepMind publishes Gemini Robotics ER 2, makes it available via the Gemini Live API, and releases an official safety tech report alongside the ASIMOV-Agentic v2 benchmark site. [3][1][7][4]
- 2026-07-30: Researcher Anirudha Majumdar publicly promotes ASIMOV-Agentic as covering unsafe task refusal, uncertainty quantification, and human escalation. [8]
- 2026-08-02: Social media amplification of Gemini Robotics 2 continues; Grok notes no fully independent third-party physical benchmark has yet verified zero-shot transfer to arbitrary platforms. [14][10]
Perspectives
Google DeepMind
Presents Gemini Robotics 2.0 as a step toward 'physical AGI' — a generalist robot capable of arbitrary human-directed tasks — with a layered stack, a formal safety tech report, and the ASIMOV-Agentic v2 benchmark.
Evolution: Consistent with prior Gemini Robotics framing; this release adds whole-body control, multi-robot heterogeneous collaboration, and a more developed safety evaluation framework at v2.
Anirudha Majumdar (researcher, ASIMOV-Agentic)
Argues ASIMOV-Agentic addresses the core requirements for agentic robot safety: refusing unsafe tasks, quantifying uncertainty, and escalating to humans when appropriate.
Evolution: First named academic voice in this thread to publicly promote ASIMOV-Agentic; provides researcher-level legitimacy beyond DeepMind's institutional communications.
Black Forest Labs
Argues for a unified multimodal foundation model covering image, video, audio, and action in a single architecture, with robotics as a downstream application of generative media training rather than a separate stack.
Evolution: New entrant to physical AI; FLUX 3 extends the company's prior image generation work into video and robotics for the first time.
The Neuron (Grant Harvey)
Bullish on the convergence thesis: training models to generate convincing physical-world video and training robots to act share underlying causal requirements, making generative video a viable route to physical intelligence.
Evolution: Consistent newsletter framing applied to both FLUX 3 and the broader wave of physical AI releases.
Ars Technica (Ryan Whitwam)
Reports DeepMind's capability claims straightforwardly, relaying the 'physical AGI' framing without significant skepticism or independent technical evaluation.
Evolution: Consistent with standard tech reporting on major AI announcements; no editorial pushback on capability claims.
Grok (xAI)
Notes that no fully independent third-party physical benchmark has yet demonstrated reliable zero-shot transfer of Gemini Robotics 2 to arbitrary platforms.
Evolution: First appearance in this thread; provides an AI-sourced articulation of the independent verification gap.
Tensions
- DeepMind builds specialized, layered robotics models (reasoning, VLA, on-device as separate tiers); Black Forest Labs argues a single unified multimodal model trained on video and action is sufficient and more efficient. [3][2][9]
- DeepMind makes ER 2 publicly available but keeps VLA and On-Device 2 behind early-access restrictions, while Grok notes no independent third-party benchmark has verified zero-shot transfer to arbitrary platforms — limiting external scrutiny of the most capable models. [2][1][10]
- FLUX 3's performance claims rest on self-reported internal evaluations against Runway and Luma, while DeepMind publishes structured benchmark results backed by peer-reviewed papers, making direct cross-system comparison unavailable. [9][3][5][6]
- DeepMind frames ASIMOV-Agentic as a general safety evaluation standard; whether other labs adopt it or define safety on their own terms remains unresolved. [8][4][5]
Sources
- [1] Google reveals Gemini Robotics 2.0, promising improved dexterity and safety — Ars Technica AI (2026-07-30)
- [2] Gemini Robotics 2 brings whole body intelligence to robots — DeepMind Blog (2026-07-28)
- [3] Gemini Robotics ER 2: powering robotics with video understanding, task orchestration, and multi-robot collaboration — DeepMind Blog (2026-07-30)
- [4] ASIMOV Benchmark v2 — reactive:gemini-robotics-2-launch
- [5] Generating Robot Constitutions & Benchmarks for Semantic Safety — reactive:gemini-robotics-2-launch
- [6] Generating Robot Constitutions & Benchmarks for Semantic Safety — reactive:gemini-robotics-2-launch
- [7] Gemini Robotics 2: Safety Evaluations - Googleapis.com — reactive:gemini-robotics-2-launch
- [8] Anirudha Majumdar on X: "Gemini Robotics 2 comes with a new open benchmark for agentic safety reasoning (ASIMOV-Agentic): refusing unsafe tasks, quantifying uncertainty, and asking for human help. Check out the safety tech report for more: https://t.co/9WMBE73epJ" / X — reactive:gemini-robotics-2-launch
- [9] 😸 Multimodal AI just got real — The Neuron (2026-07-27)
- [10] No fully independent third-party physical benchmark yet shows reliable zero-shot transfer of Gemini Robotics 2 to arbitr... — reactive:gemini-robotics-2-launch (2026-07-30)
- [11] FLUX 3 Launches: Black Forest Labs Enters Video, Audio, and Physical AI in One Model — reactive:gemini-robotics-2-launch
- [12] FLUX 3 x mimic: The Next Generation of Video-Action Models | Black Forest Labs — reactive:gemini-robotics-2-launch
- [13] Black Forest Labs Unveils First Model for Robotics in Shift to Physical AI — reactive:gemini-robotics-2-launch
- [14] Google DeepMind robots achieve whole body movement — reactive:gemini-robotics-2-launch (2026-08-02)