Skip to content
KO EN
AI 기술 Upcoming

Google DeepMind Unveils Gemini Robotics 2: Whole-Body Control, Dexterity, and Multi-Robot Collaboration

Google DeepMind has released Gemini Robotics 2 , a suite of three AI models that marks a significant leap in robotic intelligence. The new stack moves beyo

Google DeepMind has released Gemini Robotics 2, a suite of three AI models that marks a significant leap in robotic intelligence. The new stack moves beyond table-top manipulation to enable whole-body control, five-finger dexterity, and multi-robot teamwork. This release targets the core limitations of today’s robots—pre-programmed routines, inability to adapt to unpredictable environments, and poor skill transfer across different robot bodies—by introducing a layered, modular intelligence architecture.

What Happened: Three Models with a Clear Division of Labor

Gemini Robotics 2 consists of three distinct models: a vision-language-action (VLA) model, an embodied reasoning (ER) model, and an on-device VLA model. The VLA model converts visual and language input directly into motor commands, controlling humanoid robots from feet to fingertips. The ER model, based on Gemini 3.5 Flash, acts as the high-level brain—it communicates with humans, understands the physical world, and plans multi-step tasks lasting several minutes, with a context window of up to 128K tokens. The on-device VLA model is a lightweight version optimized for local execution on the robot, built on Gemini Robotics 1.5 technology and Google’s on-device Gemma models. This division of labor allows the ER model to plan and track tasks, then call the VLA model as a tool for motor execution—a design reminiscent of the human cerebrum and cerebellum.

Why It Matters: Whole-Body Control and Dexterity Breakthroughs

Previous Gemini Robotics models controlled only the upper body for table-top tasks. Gemini Robotics 2 extends control to whole-body motion for the first time. In a demonstration with Apptronik’s Apollo 2 humanoid, the robot walked to a table, picked up a watering can, and placed it into a bin on a bottom shelf—a task requiring coordinated leg and arm movement. The same model checkpoint also controls a five-fingered, 22-degree-of-freedom SharpaWave hand for tasks like tying knots and sealing ziplock bags, as well as standard two-fingered parallel grippers on a Franka Duo platform for precise insertion and packing. Success rates vary widely: multi-finger dexterity ranges from 32% (dustpan use) to 92% (unscrewing a bulb), while gripper tasks achieve 74% to 90%. Google DeepMind acknowledges that movement speed and fine dexterity still need improvement.

Our Analysis: A Layered Intelligence Strategy with Ecosystem Implications

XPLAIN AI interprets this release as a strategic move toward modular, hierarchical robot intelligence rather than a monolithic AI. By separating planning (ER) from execution (VLA), Google DeepMind enables each layer to evolve independently and allows developers to integrate custom low-level controllers. The decision to offer the ER model as a public preview while gating the VLA and on-device models suggests an intent to maintain control over core technology while fostering an ecosystem around the high-level reasoning layer. This could drive demand for cloud AI services and specialized hardware, as real-time performance may require tight integration with Google’s infrastructure. The approach mirrors the evolution of operating systems: a standardized kernel (ER) with modular drivers (VLA).

Opportunities and Risks: What Investors Should Watch

The technology could accelerate commercialization of humanoid robots with whole-body control and dexterity, benefiting hardware makers and component suppliers in motors, sensors, actuators, and batteries. Companies like Apptronik and SharpaWave may see increased demand for their platforms. Conversely, traditional industrial robot makers reliant on pre-programmed routines may face pressure to upgrade their software stacks. Google DeepMind’s gated access to the VLA models could also create a moat, potentially limiting competition but raising dependency on Google’s ecosystem. The safety benchmark ASIMOV-Agentic, released on Hugging Face under CC-BY-4.0, indicates a proactive approach to safety, but real-world validation remains early.

Uncertainties and Counter-Scenarios

Several challenges remain. First, the wide variance in dexterity success rates (32% to 92%) highlights that fine manipulation is not yet production-ready. Second, movement speed is still too slow for many industrial applications. Third, the safety benchmark is new and untested in real environments. Fourth, competition is fierce—companies like Tesla, Figure AI, and Boston Dynamics are developing similar capabilities. If Google DeepMind’s gating strategy limits adoption, or if competitors achieve comparable performance with more open access, the ecosystem could fragment. Investors should not assume immediate commercial impact; the technology is still maturing.

Key Metrics to Track

  • Developer community response and real-world use cases for the ER model public preview.
  • Improvement in VLA model success rates, especially for dexterity tasks below 50%.
  • Commercialization terms for the gated VLA and on-device models.
  • Timeline and performance of competing technologies from Tesla, Figure AI, and others.
  • Production plans and orders for humanoid hardware from Apptronik and SharpaWave.

#GoogleDeepMind #GeminiRobotics2 #HumanoidRobots #WholeBodyControl #DexterousAI #RobotIntelligence #AIInvesting #IndustrialRobots

Sources

Written by: XPLAIN AI Editorial Team · Reviewed by: XPLAIN AI Editorial Desk
This content was drafted with AI assistance based on publicly available sources and reviewed under XPLAIN AI's editorial standards.

Found an error? Request a correction →