Back to Newsroom

Gemini Robotics ER 2: powering robotics with video understanding, task orchestration, and multi-robot collaboration

By Modelverse Editorial·July 30, 2026·2 min read
Gemini Robotics ER 2: powering robotics with video understanding, task orchestration, and multi-robot collaboration

Google has launched Gemini Robotics ER 2, a significant advancement designed to serve as a sophisticated "high-level brain" for robots. This new model aims to enhance robots' utility and safety in the physical world by integrating advanced video understanding, intelligent task orchestration, and multi-robot collaboration capabilities. It represents a substantial leap forward, enabling robots to engage in real-time spatial reasoning and plan complex, multi-step actions.

Gemini Robotics ER 2 operates by allowing robots to continuously process video feeds, enabling them to monitor their own progress, identify and correct mistakes, and adapt dynamically to changing environments. While it handles the strategic, high-level planning, it seamlessly hands off the actual motor execution to lower-level Vision-Language-Action (VLA) models. A key feature is its ability to natively call external tools, such as Google Search or user-defined functions, allowing robots to gather information or perform specialized actions. Furthermore, it introduces multi-robot collaboration, empowering multiple robots to work together on intricate workflows that would be impossible for a single unit.

This powerful embodied reasoning model is now accessible to developers and researchers through the Gemini API, Google AI Studio, and the Gemini Enterprise Agent Platform. By providing a framework for orchestrating complex tasks and enabling self-correction, Gemini Robotics ER 2 empowers the creation of more adaptable, intelligent, and helpful physical AI agents. Developers can easily integrate existing low-level control interfaces as tools, streamlining the development of robots capable of tackling novel and challenging real-world scenarios.

ai-newsbreakinggoogle-deepmind

Footnotes & Primary References

Related content

A fundamental flaw leaves LLMs strikingly vulnerable to attack

It is impossible to make large language models fully secure against hacks because of a fundamental flaw in how they work, a team of researchers argue in a paper presented at the In...

Read article

Advancing the price-performance frontier with GPT-5.6

Explore lower GPT‑5.6 pricing for Luna and Terra—and how OpenAI’s more efficient models help enterprises deploy AI workflows at scale.

Read article

Best in Class: Stream PC Games and Study on the Same Laptop With GeForce NOW

Back to school means balancing assignments, deadlines and downtime. GeForce NOW makes it easy to have it all. With cloud gaming, everyday laptops used for class can also become GeF...

Read article
© 2026 Modelverse®. All rights reserved.Modelverse Newsroom