Google DeepMind released Gemini Robotics 2 on Thursday, its most advanced vision-language-action model family yet, giving humanoid robots whole-body control from feet to fingertips, dexterous manipulation on both grippers and multi-finger hands, and the ability to coordinate as a team on multi-step tasks. The launch marks a step change from the upper-body-only Gemini Robotics 1.5 and lands the intelligence layer inside Apptronik's Apollo 2 humanoid, Boston Dynamics platforms and Agile Robots hardware.
From table-top to whole-body
Gemini Robotics 2 controls the entire humanoid, translating a natural-language instruction such as "put the watering can into the green bin in the bottom shelf" into a coordinated sequence of walking, bending, picking and placing. On the Apollo 2 with SharpaWave five-finger hands, DeepMind reports 92% success on light-bulb unscrewing, 44% on tying trash bags, and 40% on sealing ziplock bags — still imperfect, but a leap for multi-finger manipulation.
Three models, one intelligence layer
The release ships as three connected models: the flagship VLA that produces motor control, Gemini Robotics ER 2 (a vision-language planner that "acts as the robot's brain" and now supports multi-robot collaboration) and Gemini Robotics On-Device 2, which adapts to a brand-new bi-arm embodiment in a few hours with fewer than 200 demonstrations. DeepMind also introduced ASIMOV-Agentic, a benchmark for embodied safety and uncertainty resolution, and says ER 2 is its safest robotics model yet in proximity and constraint-following tests.
Availability and industry context
Gemini Robotics ER 2 is now available in Google AI Studio and in private preview on the Gemini Enterprise Agent Platform, while the VLA and On-Device 2 models are gated behind an early-access partner program. The release lands weeks after Hyundai's Chung vowed an open robot reference platform with NVIDIA, alongside Generalist's GEN-1 embodied foundation model, and inside a broader physical-AI race that already includes Physical Intelligence, Skild AI and Black Forest Labs' FLUX-mimic video-action model.
Reporting based on coverage from Google DeepMind, Bloomberg, The Next Web and OODAloop.