Skip to main content
Embodied AIExplainerAug 12, 2026, 1:49 AM· 6 min read

Google DeepMind Launches Gemini Robotics 2, Bringing 'Whole-Body Intelligence' to Physical AI Models

Google DeepMind has introduced Gemini Robotics 2, a suite of AI models that allows humanoid robots to coordinate their entire bodies—from feet to fingertips—under a single learned policy. The release marks a shift from rigid, pre-programmed robotics toward adaptable systems capable of reasoning through complex physical tasks.

By Karim Mansour

Generalist AI Proponents 45%Hardware Integration Specialists 30%Pragmatic Skeptics 25%
Generalist AI Proponents
Argue that a single learned policy for whole-body control is the only scalable path to useful robotics, replacing brittle task-specific code.
Hardware Integration Specialists
Emphasize that software intelligence must still overcome physical hardware limitations, such as actuator precision and battery life.
Pragmatic Skeptics
Focus on the gap between curated demo reels and real-world reliability, pointing to low success rates in dynamic, friction-heavy tasks.

Summary

  • Google DeepMind launched Gemini Robotics 2, a suite of AI models designed to give robots 'whole-body intelligence.'
  • The system controls a humanoid robot from feet to fingertips, managing balance, crouching, and complex hand dexterity simultaneously.
  • The architecture includes three models: a high-level reasoning brain, a vision-language-action model, and a lightweight on-device version.
  • DeepMind is positioning the software as a universal intelligence layer for third-party hardware, rather than building its own robots.
  • While capable of delicate tasks like tying knots, the system still struggles with dynamic, friction-heavy actions like sweeping.

On July 30, 2026, an Apptronik Apollo 2 humanoid robot received a spoken instruction to clean up a cluttered room. For the first time, a single artificial intelligence model calculated how to bend the machine's knees, shift its center of gravity, reach out with a five-fingered hand, and grasp a watering can. This demonstration marked the launch of Gemini Robotics 2, Google DeepMind's latest suite of physical AI models. The release represents a fundamental shift in how machines interact with the physical world, moving away from rigid, pre-programmed routines toward systems that can reason about their environment and adapt on the fly.[1][2][3][6]

To understand the significance of this development, it helps to look at how robots have traditionally been operated. Most industrial robots are specialists, meticulously coded to perform a single repetitive motion, such as welding a car door or moving a box along a fixed conveyor path. If a box is dropped slightly out of place, the robot fails. Even the viral videos of humanoid robots performing backflips or parkour are typically the result of highly tuned, narrow programming designed for a specific sequence of physical feats.[1][6]

DeepMind's previous generation of robotics models began to change this by introducing vision-language-action capabilities, but they were limited to controlling a robot's upper body for tabletop tasks. The legs, core, and balance systems were essentially along for the ride, requiring humans to steady or guide the machine. Gemini Robotics 2 bridges this gap by introducing what researchers call "whole-body intelligence."[1][5]

Whole-body control is a profound mathematical and physical challenge. When a human crouches to pick up a heavy object, they instinctively adjust their foot placement, stiffen their core, and calculate the necessary grip strength to maintain balance. Translating this instinct into robotics requires a unified system that manages the entire body simultaneously. Gemini Robotics 2 achieves this through end-to-end learning, meaning a single learned policy processes camera inputs and spoken commands, and directly outputs motor controls for the legs, torso, arms, and multi-finger hands.[1][4][5]

Whole-body intelligence coordinates a robot's legs, core, and arms simultaneously to maintain balance.
Whole-body intelligence coordinates a robot's legs, core, and arms simultaneously to maintain balance.

The architecture driving this capability is not a single monolith, but a suite of three distinct models working in tandem. The core engine is the Gemini Robotics 2 vision-language-action model, which handles the split-second translation of intent into physical movement. This model is hardware-agnostic; it was demonstrated controlling both the full Apollo 2 humanoid and a Franka Duo bi-arm setup equipped with standard parallel grippers.[1][3][5]

Sitting above the vision-language-action model is Gemini Robotics ER 2, an embodied reasoning model built on the Gemini 3.5 Flash architecture. DeepMind describes this as the robot's high-level brain. ER 2 is responsible for processing complex, multi-step user instructions, observing the room, and formulating a plan. It can hold a conversation with a human operator, track its own progress through a task, and even coordinate the actions of multiple robots working in the same shared space.[1][2][3][4]

Sitting above the vision-language-action model is Gemini Robotics ER 2, an embodied reasoning model built on the Gemini 3.5 Flash architecture.

The third component addresses one of the most significant bottlenecks in modern robotics: latency and connectivity. Gemini Robotics On-Device 2 is a lightweight version of the system optimized to run locally on the robot's own hardware, without requiring a continuous connection to cloud servers. Built on Google's Gemma models, this on-device system can adapt to an entirely unfamiliar robot body in just a few hours using fewer than 200 example demonstrations.[2][4]

The Gemini Robotics 2 suite separates high-level reasoning from split-second motor control.
The Gemini Robotics 2 suite separates high-level reasoning from split-second motor control.

Google's strategic approach with this release clarifies its position in the broader robotics industry. The company is not attempting to win the hardware race by building its own humanoid chassis. Instead, DeepMind is positioning Gemini Robotics as the universal intelligence layer for machines built by third-party manufacturers like Apptronik, Boston Dynamics, and Agile Robots. This mirrors the dynamic of the early smartphone era, providing a standardized operating system for diverse hardware ecosystems.[1][4]

Despite the impressive demonstrations, the technology remains in its early stages, and its reliability varies significantly depending on the complexity of the task. Published success rates for the Apollo 2 robot using the new model show a stark contrast in capabilities. While the system achieved a 92 percent success rate for unscrewing a light bulb, its performance dropped to 36 percent when attempting to screw the bulb back in, and plummeted to 32 percent for sweeping debris into a dustpan. These figures highlight the ongoing difficulty of mastering dynamic, friction-heavy interactions.[4]

To address the inherent risks of deploying autonomous, heavy machinery in human environments, DeepMind introduced a new safety benchmark alongside the models, dubbed ASIMOV-Agentic. This framework specifically tests an agent's ability to navigate uncertainty and orchestrate safe behavior. It measures whether the embodied reasoning model can successfully refuse an unsafe tool call, or proactively halt a task and request human intervention when it predicts a high likelihood of failure.[1][2]

A major focus of the new model is fine motor control, an area where robots have historically struggled. Gemini Robotics 2 was trained to operate the 22-degree-of-freedom SharpaWave hand, allowing the Apollo 2 robot to perform delicate maneuvers such as tying knots or sealing a ziplock bag. This level of dexterity is crucial for tasks in healthcare or domestic settings, where objects are fragile and environments are not standardized for machine manipulation.[1][2][6]

Advanced dexterity models allow robots to manipulate fragile objects and perform tasks like tying knots.
Advanced dexterity models allow robots to manipulate fragile objects and perform tasks like tying knots.

Beyond individual capabilities, the embodied reasoning model introduces multi-robot collaboration. Because ER 2 can process a massive context window of interleaved text, audio, and video, it can act as a central dispatcher for a fleet of machines. If a task is too complex or heavy for a single humanoid, the reasoning model can divide the labor, directing one robot to hold an object steady while another performs the necessary manipulation.[1][3][4]

The introduction of whole-body intelligence marks a critical threshold for embodied AI. By enabling robots to reason through their movements and adapt to unpredictable environments, the technology moves the industry closer to deploying general-purpose assistants in hospitals, warehouses, and eventually homes. While the hardware and reliability metrics must still mature, the software foundation for autonomous, adaptable physical action is now firmly in place.[1][6]

Ultimately, the transition from specialized programming to generalized, whole-body AI represents the most significant bottleneck in modern robotics. As models like Gemini Robotics 2 continue to scale, the focus will increasingly shift from whether a robot can perform a task, to how reliably and safely it can execute that task in the chaotic, unstructured environments of the real world.

Definitions

Vision-Language-Action (VLA) model
An AI system that processes visual data from cameras and text or spoken commands, and translates them directly into physical motor controls.
Embodied Reasoning
The ability of an artificial intelligence to understand its physical surroundings, plan multi-step tasks, and track its progress in the real world.
End-to-end learning
A machine learning method where a single model handles the entire process from raw input, like camera pixels, to final output, like joint movement, without intermediate programming.
Whole-body control
The simultaneous coordination of a robot's entire physical structure—including legs, core, and arms—to maintain balance while executing a task.
ASIMOV-Agentic
A safety benchmark designed to test whether an AI agent can recognize unsafe actions, refuse dangerous commands, or ask for human help when uncertain.

Chronology

  1. 2025

    DeepMind releases earlier vision-language-action models capable of controlling a robot's upper body for tabletop tasks.

  2. July 30, 2026

    Google DeepMind officially announces Gemini Robotics 2, demonstrating whole-body control on the Apptronik Apollo 2 humanoid.

Analysis by camp

The Generalist AI View

The belief that scaling a single AI model across different robot bodies is the most efficient path forward.

Proponents of this approach argue that the traditional method of robotics—writing bespoke code for every specific movement—is a dead end for real-world deployment. By training a single Vision-Language-Action model that can understand its environment and adapt its policy to different hardware, developers can achieve economies of scale. This perspective views the robot chassis as a secondary commodity, with the true value lying in the universal intelligence layer that drives it.

The Hardware Integration View

The perspective that software intelligence is only half the battle, and physical mechanics remain a critical bottleneck.

Hardware specialists caution against viewing AI as a panacea for robotics. While a model might perfectly calculate the necessary center of gravity and grip strength, the physical robot must still possess the battery capacity, actuator speed, and sensor fidelity to execute the command. This camp argues that true whole-body intelligence requires a deep, physics-native co-design of both the AI and the mechanical hardware, rather than treating the robot as a simple output device for a software brain.

The Pragmatic Skeptic View

A focus on the current unreliability of these models in edge cases and complex physical interactions.

Skeptics point to the stark contrast in success rates across different tasks as evidence that these systems are not yet ready for unstructured environments. While unscrewing a lightbulb may achieve a 92 percent success rate, tasks involving complex friction and dynamic environments—like sweeping debris—often fail more than two-thirds of the time. This viewpoint stresses that until models can reliably handle the chaotic, unpredictable nature of human spaces without constant human intervention, they remain impressive laboratory demonstrations rather than deployable products.

Questions & answers

What is Gemini Robotics 2?

It is a suite of artificial intelligence models developed by Google DeepMind designed to control the physical movements of robots, translating visual and spoken commands into real-world action.

What does 'whole-body intelligence' mean?

Unlike previous models that only controlled a robot's arms, whole-body intelligence allows the AI to coordinate the legs, torso, arms, and fingers simultaneously to maintain balance and perform complex tasks.

Does Google build the robots themselves?

No. Google DeepMind is building the software intelligence layer, which is designed to be integrated into hardware built by third-party robotics companies like Apptronik.

Can these robots operate without an internet connection?

Yes. The Gemini Robotics On-Device 2 model is specifically optimized to run locally on the robot's internal hardware, ensuring functionality even without cloud connectivity.

Limits of the evidence

  • How quickly the on-device models will overcome the battery and compute constraints of current humanoid hardware.
  • Whether the low success rates on complex tasks like sweeping can be improved purely through software scaling, or if they require new tactile sensors.
  • The exact timeline for when these models will be deployed in commercial, unstructured environments like homes or hospitals.

Significance

By shifting robot control from rigid, task-specific programming to adaptable, general-purpose AI, this technology dramatically lowers the barrier to deploying robots in unpredictable human environments like hospitals, warehouses, and homes. It signals a transition where machines learn to move by understanding their surroundings rather than following a fixed script.

Sources

Source coverage

6 outlets

3 viewpoints surfaced

Generalist AI Proponents 45%Hardware Integration Specialists 30%Pragmatic Skeptics 25%
  1. [1]Google DeepMindGeneralist AI Proponents

    Gemini Robotics 2 brings whole body intelligence to robots

    Read on Google DeepMind
  2. [2]The Robot ReportHardware Integration Specialists

    Google DeepMind says Gemini Robotics 2 enables full body control

    Read on The Robot Report
  3. [3]Robotics 247Hardware Integration Specialists

    Google DeepMind announces Gemini Robotics 2 with whole-body intelligence

    Read on Robotics 247
  4. [4]RobozapsPragmatic Skeptics

    Google DeepMind released Gemini Robotics 2

    Read on Robozaps
  5. [5]Humanoid GuidePragmatic Skeptics

    Gemini Robotics 2 controls Apollo 2 whole body tasks

    Read on Humanoid Guide
  6. [6]MindStudioGeneralist AI Proponents

    What is Gemini Robotics 2?

    Read on MindStudio

Comments

Stay informed

Every angle. Every day.

Get ai stories with full source coverage and perspective breakdowns delivered to your inbox.