Google DeepMind has unveiled Gemini Robotics 2, a significant intelligence layer for next-generation robots. This release dramatically expands robotic capabilities beyond simple tabletop tasks, introducing comprehensive whole-body control, intricate five-finger dexterity, and advanced multi-robot collaboration. Delivered as three distinct models, Gemini Robotics 2 aims to overcome current limitations in robotic adaptability and skill transfer within dynamic, unpredictable environments.
The system operates through a specialized, modular design crucial for developers. Gemini Robotics 2 acts as the core Vision-Language-Action (VLA) model, translating multimodal inputs into precise motor commands for full humanoids and bi-arm robots, enabling dexterous manipulation. For high-level strategic planning, the Gemini Robotics ER 2 (Embodied Reasoning) model serves as the robot's brain, understanding human instructions and the physical world to plan complex, multi-step tasks. Built on Gemini 3.5 Flash, it processes extensive multimodal data. Complementing these, Gemini Robotics On-Device 2 provides an optimized VLA for efficient local execution. This division allows ER 2 to orchestrate missions, delegating execution to VLA models registered as callable tools, offering flexible system design.
This innovative architecture empowers robots to move beyond rigid programming into adaptive, full-body control and complex reasoning. For researchers, it paves the way for truly autonomous and versatile robotic applications, unlocking new possibilities in human-robot interaction and real-world task execution. Gemini Robotics 2 represents a pivotal step towards more capable and responsive robotic systems.
