For robots to help people in on a regular basis environments, correct spatial reasoning shouldn’t be sufficient. Robots should additionally suppose quick, timing their choices and reasoning with the real-time pace of the bodily world.
That’s why right now we’re launching Gemini Robotics ER 2, our most succesful “embodied reasoning” mannequin for robotics. Consider Gemini Robotics ER 2 as a high-level mind for robots. It permits robots to talk with people, perceive the bodily world, and plan multi-step duties. It then palms off motor execution to any given decrease stage vision-language-action (VLA) mannequin. Gemini Robotics ER 2 may also natively name instruments like Google Search to seek out info, or some other user-defined operate. The design of Gemini Robotics ER 2 permits the robotic to “suppose” about what comes subsequent whereas concurrently performing its actions.
Gemini Robotics ER 2 represents a major improve over Gemini Robotics ER 1.6. By watching steady video feeds, robots can now observe their very own progress, adapt if one thing goes unsuitable, and know precisely when to maneuver on to the subsequent step. We’re additionally introducing multi-robot collaboration, enabling robots to work collectively in shared areas and full advanced workflows a single robotic couldn’t do alone.
Gemini Robotics ER 2 is now publicly obtainable to builders by way of the Gemini API, Google AI Studio, and in non-public preview on Gemini Enterprise Agent Platform. That can assist you get began, we’re sharing examples of methods to configure the mannequin and immediate it to energy extra helpful bodily AI duties.
Advancing bodily agentic capabilities
Most duties within the bodily world are advanced and require a number of steps to finish. Gemini Robotics ER 2 is a bodily agent, orchestrating steps for the robotic and enabling it to self-correct, and generalize to extra novel conditions. To construct an agentic setup, builders can declare low-level management interfaces — like Imaginative and prescient-Language-Motion (VLA) fashions or navigation APIs — as instruments, and stream multimodal video, audio, or textual content immediately into the mannequin.
Gemini Robotics ER 2 improves this software orchestration workflow. We will consider its efficiency with robots in simulation, utilizing real-world robotic management, and even pair it with a human controlling the robotic remotely.






