Google DeepMind Launches Gemini Robotics 2 and ER 2 in Wave of Physical AI Releases
What
Google DeepMind released Gemini Robotics 2 on July 28, 2026 [1], followed by Gemini Robotics ER 2 on July 30 [2], expanding its physical AI stack with full-body humanoid control, on-device model adaptation requiring fewer than 200 examples, and real-time multi-robot collaboration. Separately, Black Forest Labs launched FLUX 3 around July 25 — a single foundation model trained across images, video, audio, and robotic actions — which powers FLUX-mimic, a robotics system tested on Audi factory assembly tasks [3]. The two releases represent distinct architectural bets: DeepMind's layered, specialized stack versus Black Forest Labs' unified multimodal-to-action model.
Why it matters
Physical AI is moving from controlled demos toward deployment-ready systems, with DeepMind claiming production-grade latency (sub-second via Gemini Live API) and on-device operation without network connectivity. The ASIMOV-Agentic benchmark DeepMind introduced could become a reference point for how the industry evaluates robot safety — specifically whether AI can refuse unsafe commands and seek human intervention under uncertainty.
Open questions
When will Gemini Robotics VLA and On-Device 2 move beyond early-access partners to general availability? [1]
Will the ASIMOV-Agentic safety benchmark gain adoption beyond DeepMind's own models as an industry-standard evaluation? [2][1]
Can FLUX 3's claim that generative video training builds physical world understanding hold up beyond self-reported comparisons against Runway and Luma? [3]
How do DeepMind's specialized robotics models and FLUX 3's unified approach compare on independent, third-party manipulation benchmarks?
Narrative
Google DeepMind released two interconnected physical AI models in late July 2026. Gemini Robotics 2, announced July 28, provides full-body humanoid control from feet to fingertips, enabling robots to walk, crouch, and manipulate objects in unstructured environments [1]. Its companion, Gemini Robotics On-Device 2, is designed for deployment on hardware with limited connectivity: it can adapt to entirely new robot morphologies in a few hours using fewer than 200 demonstration examples, running locally without network access [1]. Two days later, DeepMind published details on Gemini Robotics ER 2, the embodied reasoning layer of the stack. ER 2 achieves 91.3% accuracy on moment-finding tasks at 4x the execution speed of competing larger models and a fraction of their compute cost, and 57.4% accuracy on continuous five-level progress classification [2]. It integrates with the Gemini Live API for sub-second latency, enabling real-time orchestration of physical tasks without stop-and-think pauses [2].
A notable feature across both announcements is multi-robot collaboration: heterogeneous robots — humanoids alongside bi-arm platforms — can communicate via shared semantic understanding to hand off subtasks no single robot could complete alone [2][1]. DeepMind also introduced the ASIMOV-Agentic benchmark, which tests whether foundation models acting as high-level orchestrators can refuse unsafe tool calls from lower-level action models and proactively seek human clarification when physical feasibility is uncertain [2][1]. Access is tiered: ER 2 is publicly available via the Gemini API and Google AI Studio, while the VLA and On-Device 2 models remain limited to early-access partners [1].
Black Forest Labs took a different approach with FLUX 3, launched around July 25. Rather than a specialized robotics stack, FLUX 3 is a single foundation model trained jointly across images, video, audio, and robotic actions, capable of generating video with native audio up to 20 seconds from text or keyframes [3]. The model backbone powers FLUX-mimic, a robotics system demonstrated on Audi production line tasks involving cables, seals, and deformable parts that rule-based robots handle poorly [3]. Black Forest Labs reported that FLUX 3 outperformed Runway Gen-4.5 in 77% of comparisons and Luma Ray 3.2 in 93%, though these are self-reported internal evaluations [3]. The underlying argument from FLUX 3's proponents is architectural: training a model to generate convincing video of physical tasks requires the model to internalize causal rules about how the physical world behaves, making generative video training and robot action training mutually reinforcing [3].
Timeline
- 2026-07-25: Black Forest Labs launches FLUX 3, a unified model spanning images, video, audio, and robotic actions, with an Audi factory robotics demo. [4][5]
- 2026-07-27: The Neuron covers FLUX 3, framing the convergence of generative video and robotics training as a shared architectural foundation for physical intelligence. [3]
- 2026-07-28: Google DeepMind publishes Gemini Robotics 2, covering full-body humanoid control, on-device adaptation under 200 examples, and multi-robot collaboration. [1]
- 2026-07-30: Google DeepMind publishes Gemini Robotics ER 2, reporting 91.3% moment-finding accuracy at 4x competing model speed and sub-second latency via the Gemini Live API. [2]
Perspectives
Google DeepMind
Presents Gemini Robotics 2 and ER 2 as milestones toward general-purpose physical AI, emphasizing a layered stack (reasoning, VLA, on-device), safety via the ASIMOV-Agentic benchmark, and tiered developer access starting with ER 2.
Evolution: Consistent with DeepMind's prior Gemini Robotics line; this release adds whole-body control, multi-robot heterogeneous collaboration, and an explicit safety evaluation framework.
Black Forest Labs
Argues for a unified multimodal foundation model covering image, video, audio, and action in a single architecture, with robotics as a downstream application of generative media training rather than a separate stack.
Evolution: New entrant to the physical AI space; FLUX 3 extends the company's prior image generation work into video and robotics for the first time.
The Neuron (Grant Harvey)
Bullish on the convergence thesis: training models to generate convincing physical-world video and training robots to act in the physical world share underlying causal requirements, making generative video a route to physical intelligence.
Evolution: Consistent newsletter framing; applies it here to both FLUX 3 and the broader wave of physical AI releases.
Tensions
- DeepMind builds specialized, layered robotics models (reasoning, VLA, on-device as separate tiers); Black Forest Labs argues a single unified multimodal model trained on video and action is sufficient and more efficient. [2][1][3]
- DeepMind makes ER 2 publicly available but keeps VLA and On-Device 2 behind early-access restrictions, limiting independent verification of its most capable physical control models [1]. [1]
- FLUX 3's performance claims rest on self-reported internal evaluations against Runway and Luma, while DeepMind publishes structured benchmark results (ASIMOV-Agentic, moment-finding accuracy), making direct cross-system comparison difficult. [3][2]
Status: active and growing
Sources
- [1] Gemini Robotics 2 brings whole body intelligence to robots — DeepMind Blog (2026-07-28)
- [2] Gemini Robotics ER 2: powering robotics with video understanding, task orchestration, and multi-robot collaboration — DeepMind Blog (2026-07-30)
- [3] 😸 Multimodal AI just got real — The Neuron (2026-07-27)
- [4] FLUX 3 Launches: Black Forest Labs Enters Video, Audio, and Physical AI in One Model — reactive:gemini-robotics-2-launch
- [5] FLUX 3 x mimic: The Next Generation of Video-Action Models | Black Forest Labs — reactive:gemini-robotics-2-launch
- [6] Mimic Robotics — reactive:gemini-robotics-2-launch