Google DeepMind’s Gemini Robotics 2 brings agent layers into whole-body control

Gemini Robotics 2 separates whole-body VLA control, Robotics-ER 2 planning and tool use, and an on-device VLA model for lower-latency robot deployment.

On July 30, 2026, Google DeepMind introduced Gemini Robotics 2, extending robotics research from seeing and describing a scene to sensing, planning, and controlling a robot’s whole body. It is not one model with one deployment target. It is a set of models divided by robot embodiment, task planning, and where inference runs, which resembles a layered agent system for physical work.

Gemini Robotics 2 VLA, or vision-language-action, is designed to control full humanoid and bi-arm robots, including tasks that require dexterous hands. Robotics-ER 2 is a vision-language model and agent for multi-step planning, tool calls, and multi-robot collaboration. Robotics On-Device 2 runs the VLA on the robot itself to reduce latency and dependence on a cloud connection.

Google describes On-Device 2 as adapting to new robot embodiments in a few hours, typically with fewer than 200 demonstration examples. That is a product description, not a universal deployment guarantee. Outcomes will depend on sensors, joints, grippers, control frequency, demonstration quality, and safety constraints. Moving inference on-device also does not remove every need for cloud-based training, monitoring, or updates.

Robotics-ER 2 is the layer most similar to an embodied agent. It can understand an environment, decompose a task, call tools, and coordinate multiple robots. Google says ER 2 is available in Google AI Studio and in private preview on the Gemini Enterprise Agent Platform, while VLA and On-Device 2 are initially available to early-access partners. Availability should be read in terms of account, region, and partner eligibility rather than universal product access.

The announcement also acknowledges that the hard problems remain. Google specifically says multi-finger dexterity is still challenging. The gap between a model describing an action and a robot reliably completing it in the physical world is large. Friction, deformable objects, occlusion, collisions, battery state, and human proximity are not problems that a language model removes by itself.

For robotics teams, the three-layer design suggests a practical architecture: ER 2 handles task understanding and coordination, VLA handles vision-to-action control, and On-Device 2 handles low-latency local reactions. The physical system still needs deterministic constraints, emergency stop, human takeover, permission boundaries, and repeatable test environments. Otherwise, agent flexibility can become unpredictable behavior.

Google provides a safety framework and safety report for Gemini Robotics 2, but the scope of those documents should not be treated as certification for every physical robot. Deployment still needs hardware-specific collision, load, failure, communication-loss, and human-safety testing, with a direct path for a person to stop or take over. Embodied AI safety is a system problem spanning models, controllers, hardware, and on-site procedures.

The larger signal is that Google is extending Gemini’s agent capabilities into environments where actions have physical consequences, not only digital tools. The first evaluation question for a business should not be whether a robot looks human. It should be whether a defined task completes reliably under constraints. Success rate, retries, latency, energy, human takeovers, and safety events will say more than one demonstration.

MODULE.002 //

More insights

Ideas on websites, AI automation, digital marketing, AI news, and VMTS updates.