Google DeepMind has released Gemini Robotics ER 2, the next generation of its embodied-reasoning model for robots, opening a public preview on Google AI Studio and the Gemini API and a private preview on the new Gemini Enterprise Agent Platform. The model is pitched as a high-level "brain" that a robot can pair with any lower-level vision-language-action (VLA) controller, letting the machine reason about the next step while its body keeps moving.
What ER 2 does differently
Gemini Robotics ER 2 focuses on multi-step planning and spatial logic rather than motor control. DeepMind highlights four capabilities: multi-robot collaboration in a shared workspace, advanced spatial logic over objects and movement, success tracking that knows when a task is complete or has to be retried, and tool use that lets the model call Google Search or Google Calendar to fill in real-world context before it acts. It also holds a conversation with the human operator and can explain what it just did.
Benchmarks: ahead of Opus 5, GPT 5.6 Sol and Gemini 3.6 Flash
On DeepMind's own evals, Gemini Robotics ER 2 hits 87.7% on image-based success detection, 82.4% on video-based success detection, 78.5% on the ERQA question-answering benchmark and 65.7% on generalised instrument reading — top of the leaderboard against Opus 5, GPT 5.6 Sol, Gemini 3.6 Flash and its predecessor Gemini Robotics ER 1.6. On progress classification it jumps to 57.4% from 42.7% for ER 1.6.
Safety takes a big step up
Safety is the number DeepMind is loudest about. ER 2 hits 97.9% on safety instruction following, up from 47.2% for ER 1.6, and 93.0% on the one-metre human-proximity test versus 51.1% for the previous model. The model is designed to detect when humans are nearby and can trigger safety tool calls to bring the robot to a controlled stop.
Building on Gemini Robotics 2
ER 2 slots into DeepMind's broader physical-AI stack alongside Gemini Robotics 2's whole-body humanoid control, and lands the same week as Anthropic's Model Hardware Standard — a signal that the frontier labs are converging on standard "brain" interfaces that any robot vendor can plug into. The model is available in Google AI Studio, the Gemini API and the Enterprise Agent Platform, with multimodal inputs (text, image, video, audio) and Live API support.
Reporting based on coverage from Google DeepMind, The Robot Report and Technobezz.