Gemini Robotics ER 2: powering robotics with video understanding, task orchestration, and multi-robot collaboration
2026-08-20 · Google DeepMind
Gemini Robotics ER 2: A New High-Level Brain for Robots
Introduction
Google has launched Gemini Robotics ER 2, its most capable embodied reasoning model to date. Positioned as a high-level brain for robots, it enables real-time spatial reasoning, multi-step task planning, natural conversation with humans, and understanding of the physical world. The model delegates low-level motor execution to any chosen Vision-Language-Action (VLA) model while maintaining the ability to think ahead.
Gemini Robotics ER 2 can natively call tools such as Google Search or user-defined functions. Its architecture is specifically designed to allow the robot to reason about the next action while simultaneously performing current movements.
Major Improvements Over ER 1.6
Compared to its predecessor, Gemini Robotics ER 2 introduces substantial upgrades. The most significant is continuous video understanding. By watching live video feeds, robots can now track their own progress, detect when something goes wrong, adapt accordingly, and know precisely when to transition to the next step.
The model also introduces native multi-robot collaboration, allowing multiple robots to work together in shared physical spaces to accomplish complex workflows that would be impossible for a single robot.
Advancing Physical Agentic Capabilities
Most real-world tasks are multi-step and complex. Gemini Robotics ER 2 functions as a physical agent that orchestrates these steps, enables self-correction, and generalizes to novel situations.
Developers can declare low-level control interfaces — such as VLA models or navigation APIs — as tools. Multimodal inputs including video, audio, and text can be streamed directly into the model. ER 2 shows clear improvements in tool orchestration performance.
Evaluations across three control modes demonstrate consistent superiority over ER 1.6:
- Real VLA control
- Simulation VLA control
- Human tele-operation
Low-Latency Integration with Gemini Live API
High-level reasoning in robotics is tightly coupled with execution speed. To address this, Gemini Robotics ER 2 integrates with the Gemini Live API through a bidirectional streaming endpoint optimized for latency-sensitive applications.
This enables fluid orchestration where the model can command action models and robotics APIs to complete multi-step tasks without jarring “stop-and-think” interruptions.
A practical demonstration uses Boston Dynamics’ Spot robot. Gemini Robotics ER 2 orchestrates Spot’s navigation and manipulator APIs, allowing the robot to fetch objects in response to natural language commands — for example, retrieving a popcorn snack. The demonstration code is publicly available on GitHub.
Unlocking Temporal Intelligence
A fundamental challenge in robotics is determining when a task has been completed to specification. Gemini Robotics ER 2 delivers a significant leap in video understanding and progress tracking.
Two core capabilities have been advanced:
Continuous Progress Classification
Progress classification allows the robot to assess how far a task has advanced. The model categorizes each video frame into one of five buckets: 0-20%, 20-40%, 40-60%, 60-80%, or 80-100%. This real-time situational awareness enables the robot to adjust actions dynamically or retry failed steps without restarting the entire workflow.
On progress classification benchmarks, Gemini Robotics ER 2 achieves 57.4% accuracy, outperforming previous generation models.
Precision Moment-Finding
Moment-finding measures the model’s ability to identify the exact frame when a critical event occurs — such as the moment a cup is full while pouring coffee. This capability allows robots to switch between tasks at the right time, verify success, and suggest corrections when necessary.
Gemini Robotics ER 2 achieves 91.3% accuracy on moment-finding tasks.
Availability
Gemini Robotics ER 2 is now publicly available to developers via the Gemini API and Google AI Studio. It is also available in private preview on the Gemini Enterprise Agent Platform. These channels enable developers to begin building more capable and helpful physical AI agents immediately.
The combination of powerful video understanding, low-latency task orchestration, and multi-robot collaboration represents a meaningful step toward robots that can be genuinely useful in everyday human environments.