Google DeepMind announced Gemini Robotics 2 on Thursday, saying the new AI robotics model can operate a humanoid robot’s legs, torso, arms and hands together rather than focusing mainly on upper-body manipulation. That matters because a robot that can only reach and grab is a lab demo with knees; a robot that can walk, crouch, stretch and handle objects has a better shot at doing useful work in messy physical spaces.
The company said the previous Gemini Robotics model centered on a humanoid robot’s upper body. Gemini Robotics 2 expands control across the machine, which Google DeepMind described as a step toward more complicated real-world tasks that require coordinated movement across the body.
Google’s demos show Apptronik’s Apollo 2 humanoid doing jobs that require more than arm movement. In videos shared by the company, Apollo 2 bends down to pick up a watering can and identifies specific items on a shelf before removing them. Google DeepMind also acknowledged that its robots still need improvement in movement speed, a useful caveat amid the usual robotics video problem: edits and controlled environments can make slow, brittle systems look more competent than they are.
What can Gemini Robotics 2 do?
According to Google DeepMind, Gemini Robotics 2 lets humanoid robots combine locomotion and object manipulation, including walking, crouching, reaching and grasping. In plain terms, the model is meant to help the robot use its whole body to complete a physical task, rather than treating the legs as a mobility platform and the arms as a separate problem.
Google DeepMind also said the model improves hand control. It can work with more complex five-fingered robot hands, which the company says enables tasks such as closing a Ziploc bag, tying a trash bag and unscrewing a lightbulb. Those are small household actions, but they are also the kind of fiddly contact-rich tasks that expose how far robot dexterity still has to go.
The company is updating Gemini Robotics ER as well. Google DeepMind describes ER, short for embodied reasoning, as a vision-language model that helps robots interpret their surroundings, follow instructions and carry out tasks with multiple steps. If a language model is the part that reasons over instructions and context, the robotics stack still has to translate that into motion without smashing into the world. For more background on the model side of that equation, Kernel has an explainer on how LLMs work when they answer a prompt.
Google DeepMind said Gemini Robotics ER 2 is better at working over longer spans of time and can recognize when a task starts and finishes. The company also said ER 2 can coordinate different kinds of robots. In one Google video, Apollo 2 directs Google’s dual-arm robot to place tools in a bin while cleaning a garage.
Safety is part of the update, at least by Google DeepMind’s account. The company called Gemini Robotics ER 2 its safest robotics model so far, saying it can better notice nearby humans, trigger safety-related tool calls and stop the robot when someone gets too close.
Google DeepMind also said it improved the Gemini Robotics on-device model, which can run locally on a robot without an internet connection. The company said that model can now adapt more quickly to new robot bodies, including machines with very different shapes, sensors and degrees of freedom.
This story draws on original reporting from The Verge.