AI Policy Wiki
Dashboard

Gemini Robotics 2

medium confidence · updated 2026-08-02

Google DeepMind's July 2026 robotics model family — a vision-language-action model with whole-body humanoid control, the Gemini Robotics ER 2 embodied-reasoning model, and Gemini Robotics On-Device 2, introduced with the ASIMOV-Agentic safety benchmark.

Gemini Robotics 2 is a family of three robotics models released by Google DeepMind on July 30, 2026: a vision-language-action model of the same name, the Gemini Robotics ER 2 embodied-reasoning model, and Gemini Robotics On-Device 2. The release extended the line from upper-body, table-top manipulation to whole-body control of humanoid robots, and added multi-robot collaboration and a safety benchmark for agentic orchestration (Source: deepmind.google).

Except where a second source is cited, every statement on this page derives from that announcement by the developer. The separate safety technical report, the ER 2 and On-Device 2 model cards, and the ASIMOV-Agentic dataset are named in the announcement but have not been retrieved, and no independent evaluation of the family has been published.

Reporting on July 31, 2026 recorded a qualification on the demonstration footage, which showed robots clearing trash, inserting a tape into a boombox, screwing in a lightbulb and tying a garbage bag. Google described the footage as real-time and the robots as fully autonomous, but said each task shown had been trained in advance using human teleoperation, video examples and simulation, and that it is not presenting the robots as general-purpose machines. Carolina Parada, head of robotics at Google DeepMind, told Wired that "the safety question is even more pressing because you're putting them in a lot of other situations" (Source: techbriefly.com).

Model family

ModelTypeFunction
Gemini Robotics 2Vision-language-action (VLA)Converts vision and language input into motor control; controls full humanoids and other bi-arm robots
Gemini Robotics ER 2Vision-language model (VLM)Embodied reasoning; communicates with humans, plans multi-step tasks lasting several minutes, coordinates the VLA
Gemini Robotics On-Device 2Vision-language-action (VLA)Runs locally on robotic devices; adapts to new embodiments

In the announced architecture, ER 2 acts as the high-level planner: it observes the environment, reasons about the steps a task requires, directs the VLA to carry out actions, and tracks progress. Google DeepMind states that this arrangement lets robots execute multi-step tasks involving hundreds of decisions, self-correct when a step fails, and generalize to new goals, and that ER 2 can identify when tasks begin and end and pinpoint when key events occur (Source: deepmind.google).

Capabilities and benchmarks

Google DeepMind reports results from a single model checkpoint controlling three embodiments: Apptronik's Apollo 2 humanoid with SharpaWave hands, Apollo 2 with Inspire hands, and the Franka Duo with a Robotiq gripper. Each figure below is a success rate averaged over multiple tasks within a skill category, except the multi-finger results, which are reported per task (Source: deepmind.google).

CategoryEmbodimentTaskSuccess rate
General whole-body manipulationApollo 2, Inspire handsPick up from shelf76.3%
General whole-body manipulationApollo 2, Inspire handsPick up from table68.4%
General whole-body manipulationApollo 2, Inspire handsPick up from floor45.7%
Gripper dexterityFranka DuoPrecise insertion tasks89.6%
Gripper dexterityFranka DuoDiverse tool kitting78.9%
Gripper dexterityFranka DuoGeneral pick and place74.2%
Multi-finger dexterityApollo 2, SharpaWave handsUnscrew bulb92%
Multi-finger dexterityApollo 2, SharpaWave handsTie trash bag44%
Multi-finger dexterityApollo 2, SharpaWave handsZiplock40%
Multi-finger dexterityApollo 2, SharpaWave handsScrew bulb36%
Multi-finger dexterityApollo 2, SharpaWave handsDustpan32%

Google DeepMind characterizes the results as "medium to high" for whole-body and gripper-based dexterous tasks and states that "multi-finger dexterous manipulation remains challenging." It also notes that the robots "have more to advance in movement speed" (Source: deepmind.google).

The SharpaWave hand used in the multi-finger evaluations is a five-fingered end effector with 22 degrees of freedom. A demonstration described in the announcement has Apollo 2 respond to the instruction "put the watering can into the green bin in the bottom shelf" by walking to a table, picking up the object, walking to the shelves, and placing it (Source: deepmind.google).

Training and architecture

Parameter counts, training compute, and data composition were not disclosed. Gemini Robotics On-Device 2 is described as natively multi-embodiment and as inheriting the "motion transfer" techniques introduced with Gemini Robotics 1.5. Google DeepMind states that it adapts to new bi-arm embodiments in a few hours, typically with fewer than 200 examples, including embodiments differing substantially in shape, sensors, and degrees of freedom, and reports demonstrations on the Dexmate, SO101, and Trossen platforms (Source: deepmind.google).

Safety and evaluations

The release introduced ASIMOV-Agentic, a benchmark for agentic safety orchestration and uncertainty resolution, published as a dataset on Hugging Face. Google DeepMind describes it as measuring an embodied-reasoning agent's ability to refuse unsafe tool calls issued by a VLA, to predict whether a task is possible, and to request human intervention when uncertain (Source: deepmind.google).

Google DeepMind states that Gemini Robotics ER 2 is its "safest robotics model to date in safety constraint following and human proximity benchmarks," and that it can detect nearby humans, trigger safety tool calls, and bring a robot to a safe stop when someone approaches too closely — a capability it identifies as a requirement in collaborative safety standards. The company published a separate Gemini Robotics 2: Safety Technical Report. These are the developer's own evaluations, and no third-party assessment has been published (Source: deepmind.google).

Availability

Gemini Robotics ER 2 is available through Google AI Studio and in private preview on the Gemini Enterprise Agent Platform. The VLA and On-Device models are restricted to early-access partners through a trusted-tester program. Model cards were published for ER 2 and for On-Device 2 (Source: deepmind.google).

Google DeepMind names Apptronik, Boston Dynamics, and Agile Robots as hardware partners on the release. Carolina Parada is the announcement's named author (Source: deepmind.google).

The line runs from Gemini Robotics (March 2025) through Gemini Robotics 1.5 (September 2025) and Gemini Robotics-ER 1.6 (April 2026) to the July 2026 release (Source: deepmind.google). The announcement does not state which Gemini generation the family is built on.

Relationships