Google has rolled out Gemini Robotics ER 2, a new model designed to help robots better understand their environment and handle complex tasks while interacting with humans. Unlike traditional systems, this model acts like a high-level planner, coordinating robot actions in real time and reducing delays between decisions.
Gemini Robotics ER 2 can analyze continuous streams of video, audio, and text, while simultaneously using external tools such as Google Search and navigation systems. This multi-tasking ability allows robots to plan their next moves even as current actions are still underway. According to Google, the model achieved 57.4% accuracy in determining task progress and 91.3% accuracy in pinpointing exact moments when specific events happen, like stopping the pouring of a liquid.
One notable demonstration featured Boston Dynamics’ Spot robot following a spoken command to locate, pick up, and return an object. The model also supports teamwork between robots. Google showed how Apptronik’s Apollo 2 and the Franka F3 Duo could split and hand off tasks based on their individual strengths. The system can detect errors, monitor instrument readings, and ensure safety by recognizing spills, slips, or when a person enters a robot’s workspace, pausing actions instantly.
Developers can access Gemini Robotics ER 2 via the Gemini API and Google AI Studio, with an exclusive preview available through the Gemini Enterprise Agent Platform. The low-latency streaming endpoint connects robots in real time, enabling faster and more responsive control. This innovation marks a step forward in robotics, blending AI reasoning with physical execution to handle dynamic, real-world environments.
This material is for informational purposes and does not constitute financial advice.



