Gemini Robotics 2 is a family of three robotics models released by Google DeepMind on July 30, 2026: a vision-language-action model of the same name, the Gemini Robotics ER 2 embodied-reasoning model, and Gemini Robotics On-Device 2. The release extended the line from upper-body, table-top manipulation to whole-body control of humanoid robots, and added multi-robot collaboration and a safety benchmark for agentic orchestration (Source: deepmind.google).
Except where a second source is cited, every statement on this page derives from that announcement by the developer. The separate safety technical report, the ER 2 and On-Device 2 model cards, and the ASIMOV-Agentic dataset are named in the announcement but have not been retrieved, and no independent evaluation of the family has been published.
Reporting on July 31, 2026 recorded a qualification on the demonstration footage, which showed robots clearing trash, inserting a tape into a boombox, screwing in a lightbulb and tying a garbage bag. Google described the footage as real-time and the robots as fully autonomous, but said each task shown had been trained in advance using human teleoperation, video examples and simulation, and that it is not presenting the robots as general-purpose machines. Carolina Parada, head of robotics at Google DeepMind, told Wired that "the safety question is even more pressing because you're putting them in a lot of other situations" (Source: techbriefly.com).
Model family
| Model | Type | Function |
|---|---|---|
| Gemini Robotics 2 | Vision-language-action (VLA) | Converts vision and language input into motor control; controls full humanoids and other bi-arm robots |
| Gemini Robotics ER 2 | Vision-language model (VLM) | Embodied reasoning; communicates with humans, plans multi-step tasks lasting several minutes, coordinates the VLA |
| Gemini Robotics On-Device 2 | Vision-language-action (VLA) | Runs locally on robotic devices; adapts to new embodiments |
In the announced architecture, ER 2 acts as the high-level planner: it observes the environment, reasons about the steps a task requires, directs the VLA to carry out actions, and tracks progress. Google DeepMind states that this arrangement lets robots execute multi-step tasks involving hundreds of decisions, self-correct when a step fails, and generalize to new goals, and that ER 2 can identify when tasks begin and end and pinpoint when key events occur (Source: deepmind.google).
Capabilities and benchmarks
Google DeepMind reports results from a single model checkpoint controlling three embodiments: Apptronik's Apollo 2 humanoid with SharpaWave hands, Apollo 2 with Inspire hands, and the Franka Duo with a Robotiq gripper. Each figure below is a success rate averaged over multiple tasks within a skill category, except the multi-finger results, which are reported per task (Source: deepmind.google).
| Category | Embodiment | Task | Success rate |
|---|---|---|---|
| General whole-body manipulation | Apollo 2, Inspire hands | Pick up from shelf | 76.3% |
| General whole-body manipulation | Apollo 2, Inspire hands | Pick up from table | 68.4% |
| General whole-body manipulation | Apollo 2, Inspire hands | Pick up from floor | 45.7% |
| Gripper dexterity | Franka Duo | Precise insertion tasks | 89.6% |
| Gripper dexterity | Franka Duo | Diverse tool kitting | 78.9% |
| Gripper dexterity | Franka Duo | General pick and place | 74.2% |
| Multi-finger dexterity | Apollo 2, SharpaWave hands | Unscrew bulb | 92% |
| Multi-finger dexterity | Apollo 2, SharpaWave hands | Tie trash bag | 44% |
| Multi-finger dexterity | Apollo 2, SharpaWave hands | Ziplock | 40% |
| Multi-finger dexterity | Apollo 2, SharpaWave hands | Screw bulb | 36% |
| Multi-finger dexterity | Apollo 2, SharpaWave hands | Dustpan | 32% |
Google DeepMind characterizes the results as "medium to high" for whole-body and gripper-based dexterous tasks and states that "multi-finger dexterous manipulation remains challenging." It also notes that the robots "have more to advance in movement speed" (Source: deepmind.google).
The SharpaWave hand used in the multi-finger evaluations is a five-fingered end effector with 22 degrees of freedom. A demonstration described in the announcement has Apollo 2 respond to the instruction "put the watering can into the green bin in the bottom shelf" by walking to a table, picking up the object, walking to the shelves, and placing it (Source: deepmind.google).
Training and architecture
Parameter counts, training compute, and data composition were not disclosed. Gemini Robotics On-Device 2 is described as natively multi-embodiment and as inheriting the "motion transfer" techniques introduced with Gemini Robotics 1.5. Google DeepMind states that it adapts to new bi-arm embodiments in a few hours, typically with fewer than 200 examples, including embodiments differing substantially in shape, sensors, and degrees of freedom, and reports demonstrations on the Dexmate, SO101, and Trossen platforms (Source: deepmind.google).
Safety and evaluations
The release introduced ASIMOV-Agentic, a benchmark for agentic safety orchestration and uncertainty resolution, published as a dataset on Hugging Face. Google DeepMind describes it as measuring an embodied-reasoning agent's ability to refuse unsafe tool calls issued by a VLA, to predict whether a task is possible, and to request human intervention when uncertain (Source: deepmind.google).
Google DeepMind states that Gemini Robotics ER 2 is its "safest robotics model to date in safety constraint following and human proximity benchmarks," and that it can detect nearby humans, trigger safety tool calls, and bring a robot to a safe stop when someone approaches too closely — a capability it identifies as a requirement in collaborative safety standards. The company published a separate Gemini Robotics 2: Safety Technical Report. These are the developer's own evaluations, and no third-party assessment has been published (Source: deepmind.google).
Availability
Gemini Robotics ER 2 is available through Google AI Studio and in private preview on the Gemini Enterprise Agent Platform. The VLA and On-Device models are restricted to early-access partners through a trusted-tester program. Model cards were published for ER 2 and for On-Device 2 (Source: deepmind.google).
Google DeepMind names Apptronik, Boston Dynamics, and Agile Robots as hardware partners on the release. Carolina Parada is the announcement's named author (Source: deepmind.google).
Related models
The line runs from Gemini Robotics (March 2025) through Gemini Robotics 1.5 (September 2025) and Gemini Robotics-ER 1.6 (April 2026) to the July 2026 release (Source: deepmind.google). The announcement does not state which Gemini generation the family is built on.