**Google DeepMind Releases Gemini Robotics 2: A New Intelligence Layer for Embodied Robots**
Google DeepMind has delivered a significant step forward for robotics with the release of **Gemini Robotics 2**, the intelligence layer powering its next-generation robots. The update marks a decisive move beyond table-top, single-task robots, enabling **whole-body control, five-finger dexterity, and multi-robot collaboration**.
—
### Key Takeaways
* **Three Models, One Vision:** Gemini Robotics 2 ships as a coordinated stack:
1. **Gemini Robotics 2 (VLA):** Converts vision and language into full-body motor control.
2. **Gemini Robotics ER 2:** The high-level “brain” powered by Gemini 3.5 Flash, capable of planning complex, multi-step tasks.
3. **Gemini Robotics On-Device 2:** An efficient, local VLA for robots that can’t rely on cloud connectivity.
* **From Tables to Worlds:** Unlike prior models limited to upper-body tabletop tasks, Gemini Robotics 2 delivers **whole-body control**. A demo with Apptronik’s Apollo 2 shows the robot walking to a table, picking up a watering can, and placing it on a high shelf.
* **Dexterity Across Bodies:** A single model checkpoint successfully controls **Apptronik’s SharpaWave five-finger hand**, **Inspire hands**, and a **Franka Duo Robotiq gripper**. Success rates vary by task, with multi-finger dexterity (32%–92%) identified as the current weak point.
* **Public Preview & Access:** The **Gemini Robotics ER 2** model is available in public preview via AI Studio and the Gemini API. The higher-tier VLA and on-device models remain gated for early-access partners.
—
### 3 Models and What They Do
**Gemini Robotics 2 (VLA)** is the core controller. It ingests vision and language to output motor commands, enabling humanoid robots to move from feet to fingertips and handle dexterous manipulation on both multi-finger and parallel grippers.
**Gemini Robotics ER 2** acts as the system’s strategic planner. Built on Gemini 3.5 Flash, it processes up to 128k tokens of multimodal input (text, image, video, audio) to execute long-horizon tasks, outputting up to 64k tokens of plans.
**Gemini Robotics On-Device 2** is designed for low-latency, offline operation. It adapts to entirely new robot bodies in hours using fewer than 200 examples, demonstrated on platforms like Dexmate and SO101.
—
### Whole-Body Control on Apptronik Apollo 2
Previous Gemini Robotics iterations were confined to the robot’s upper body. Gemini Robotics 2 changes that by enabling **full locomotion and manipulation**. Given a task like *”put the watering can into the green bin on the bottom shelf,”* Apollo 2 walks to the table, grasps the can, navigates to the shelf, and completes the placement. Google DeepMind notes that **movement speed remains a key area for advancement**.
—
### Dexterity Across Hands and Grippers
The same VLA checkpoint showcases broad adaptability:
* **Five-finger tasks** on Apollo 2 include unscrewing a bulb (92% success) and tying a trash bag (44%).
* **Two-finger gripper tasks** on Franka Duo include precise insertions (89.6%) and diverse tool kitting (78.9%).
—
### Progress Classification, Moment Finding, and Tool Orchestration
Beyond execution, ER 2 introduces advanced cognitive benchmarks:
* **Progress Classification:** The model assigns each video frame a status between 0–100%, achieving 57.4% accuracy.
* **Moment Finding:** It pinpoints the exact frame for an event (e.g., when to stop pouring), with a mean error of just 0.96 seconds.
* **Tool Orchestration:** ER 2 integrates with the Gemini Live API to chain tools and eliminate “stop-and-think” pauses, while also supporting native calls to Google Search or custom user functions.
—
### Multi-Robot Collaboration
Gemini Robotics 2 enables **heterogeneous robot teams**. A reasoning model (ER 2) on a wheeled rover can collaborate with a humanoid (Apollo 2) by sharing a common semantic understanding and handing off subtasks as needed.
—
### On-Device 2: Adapting to New Robot Bodies
Gemini Robotics On-Device 2 inherits motion-transfer technology from Gemini Robotics 1.5. It proves that **data scaling works**: comparing On-Device 2 to On-Device 1 on new platforms shows success rates jumping from single digits to over 50% with the same hardware.
—
## FAQ
**Q: What is Gemini Robotics 2?**
A: It is a vision-language-action (VLA) model that serves as the intelligence layer for advanced robots, turning camera and language inputs into precise motor actions.
**Q: Can it control a full humanoid?**
A: Yes. It can drive “full humanoids from feet to fingertips,” as demonstrated on Apptronik Apollo 2.
**Q: What hardware is it running on?**
A: The release features Apptronik Apollo 2, Franka Duo, and Boston Dynamics Spot. The on-device model has been tested on Dexmate, SO101, and Trossen.
**Q: Is it available to the public?**
A: Partially. The ER 2 model is in public preview. The VLA and on-device models are currently limited to trusted partners.
**Q: What safety measures are in place?**
A: Google DeepMind references ASIMOV-Agentic, a safety benchmark on Hugging Face, and implements human proximity monitoring and uncertainty detection.
—
## Conclusion
Gemini Robotics 2 represents a major consolidation of AI and robotics. By unifying whole-body control, five-finger dexterity, and long-horizon reasoning into a single stack, Google DeepMind is pushing robots closer to true general-purpose intelligence. While speed and out-of-distribution generalization remain challenges, the release—spanning cloud-based ER 2 and efficient on-device models—establishes a powerful new baseline for embodied AI.



