RoboticsMachines in the physical world
Gemini Robotics 2 controls whole humanoid bodies, but DeepMind's research lead says moving reliably to new robots is still unsolved
Gemini Robotics 2 pairs a publicly available reasoning model, ER 2, with partner-only action models that control entire humanoid bodies. DeepMind's Keerthana Gopalakrishnan says reliably moving to new robot bodies without extra training remains unsolved, and most performance figures are Google's own.

The system has two tiers. ER 2, a reasoning model from Gemini's Flash line, plans tasks and calls an action model the way a digital agent calls tools. Developers can use ER 2 now as an API preview. The action model, called a vision-language-action model (VLA), controls a humanoid from fingertips to feet. It and the smaller On-Device 2 model are limited to early-access partners, and lower-level stabilization still runs beneath the VLA. [1] [2] [6]
Moving to a new robot body is still the main open problem. Gopalakrishnan says there is no strong precedent for a model handling many tasks reliably on a body it was not trained on. Google says adaptation takes fewer than 200 examples, but that is Google's own figure. The only outside adaptation result in the reviewed sources is a preprint, not yet peer-reviewed. It reports that fine-tuning through Google's managed API worked well on a real humanoid. [1] [2] [4]
The evidence on performance is thin. Google's figure for whole-body manipulation success ranges from about 46% to 76%, depending on the task. Gopalakrishnan says many failures happen in the hand-off between the two models, when deciding a subtask is done and passing control on. The host's informal test took about six seconds from prompt to tool call, which sits awkwardly with Google's claim of sub-second streaming. [1] [2] [7]
Gopalakrishnan expects the first deployments in pick-and-place work and in tasks where mistakes can be retried. She treats safety around the robot as separate from whether its goals are right. A robot falling over is a hazard even when nothing is wrong with its goals, and a robot that cannot see should ask for help. Google's API docs also warn that the model can hallucinate, and they make developers responsible for safety. [1] [6]
developers can prototype robot planning through the ER 2 API today, and a preprint reports managed fine-tuning works.
Read the full assessment
Implication: teams may adapt Google's action models without owning the weights, but they face platform lock-in and reliability that has not been independently checked.
Executive brief
Google DeepMind's Gemini Robotics 2 controls an entire humanoid, from fingertips to feet, with a single learned policy (DeepMind). Even so, research lead Keerthana Gopalakrishnan says robotics is still in its "GPT-2 era." Her reason is that robots cannot yet move to a new body and run many tasks reliably without extra training on that body (transcript, 00:51:02). Developers can use the reasoning model, ER 2, through the API now. The two models that actually move robots are limited to partners. The only third-party adaptation result found is a preprint.
Read the full section
Google DeepMind's Gemini Robotics 2 controls an entire humanoid, from fingertips to feet, with a single learned policy (DeepMind). Even so, research lead Keerthana Gopalakrishnan says robotics is still in its "GPT-2 era." Her reason is that robots cannot yet move to a new body and run many tasks reliably without extra training on that body (transcript, 00:51:02). Developers can use the reasoning model, ER 2, through the API now. The two models that actually move robots are limited to partners. Most performance figures come from Google. The only third-party adaptation result found is a preprint.
What changed and event timeline
Gemini Robotics 1 launches
The first VLA in the family, focused on gripper dexterity. Gopalakrishnan uses it as the reference point when describing progress to multi-finger hands (;).
On-Device variant released
A smaller model built to run locally on robot hardware, not in the cloud ().
Gemini Robotics 2 suite
Three models ship: the VLA, ER 2 and On-Device 2. Google adds whole-body control, multi-robot collaboration and a new safety benchmark, ASIMOV-Agentic ().
CLIFT preprint
Researchers fine-tune Gemini Robotics On-Device on a real humanoid using only the managed API. They report near-perfect task success after two improvement cycles ().
Trade press coverage
The Robot Report repeats the vendor claims, including a 22-degree-of-freedom hand. It also notes DeepMind's own admission that robot movement speed still needs work ().
Podcast interview
Gopalakrishnan talks with host Nathan Labenz about the architecture, latency, cross-embodiment limits and safety ().
Capabilities and access
- ER 2 is the reasoning layer. Input is limited to 131,072 tokens and output to 65,536. ER 1.6 is being retired (Gemini API docs).
- Gemini Robotics 2 (the VLA) and On-Device 2 are open only to early-access partners who apply (DeepMind).
- Robots shown: Apptronik Apollo 2, Franka Duo, Dexmate, Trossen and SO101. Development partners are Apptronik, Boston Dynamics and Agile Robots (DeepMind).
Read the full section
- ER 2 is the reasoning layer. It is available in preview as
gemini-robotics-er-2-previewandgemini-robotics-er-2-streaming-preview. Input is limited to 131,072 tokens and output to 65,536. It supports pointing, bounding boxes, function calling and video progress tracking. ER 1.6 is being retired (Gemini API docs). - Gemini Robotics 2 (the VLA) and On-Device 2 are open only to early-access partners who apply (DeepMind).
- Robots shown: Apptronik Apollo 2, Franka Duo, Dexmate, Trossen and SO101. Development partners are Apptronik, Boston Dynamics and Agile Robots (DeepMind).
Technical analysis for researchers and developers
- ER 2 plans the task and calls a VLA as a tool, much like a digital agent calls functions.
- Labenz says it is built on Gemini 3.5 Flash. Gopalakrishnan says only "the Flash line," and the API docs name no base model.
- By her estimate, the context window holds about three minutes of robot memory. She suggests replacing older frames with text summaries (00:32:55).
Read the full section
- Two-tier design. ER 2 plans the task and calls a VLA as a tool, much like a digital agent calls functions. The VLA turns language and images into joint commands. Lower-level stabilization still sits beneath the VLA (transcript, 00:24:46, 00:49:29).
- Base model. Labenz says it is built on Gemini 3.5 Flash. Gopalakrishnan says only "the Flash line," and the API docs name no base model.
- Context. By her estimate, the context window holds about three minutes of robot memory. She suggests replacing older frames with text summaries (00:32:55).
- Where failures come from. Many failures arise in the hand-off between the two models: deciding when a subtask is finished and passing control on, which adds latency (00:51:02).
Claims and evidence
The host's latency result sits uneasily with Google's claim that streaming runs at "sub-second" speed (Google blog). The two may have used different endpoints. No independent benchmark of the VLA itself was found.
Read the full section
| Claim | Status |
| Whole-body manipulation success of 45.7–76.3%; multi-finger dexterity 36–92%, depending on task (DeepMind) | Vendor-reported |
| ER 2 moment-finding accuracy of 91.3%; progress classification of 57.4% (Google blog) | Vendor-reported |
| Adapts to a new robot with fewer than 200 examples, in "a few hours" (DeepMind) | Vendor-reported |
| API fine-tuning beats a leading open-weight VLA trained on the same demonstrations (arXiv) | Third-party preprint, not peer-reviewed |
| Base Flash and ER 2 performed the same on simple 2D tests; about 6 seconds from prompt to tool call (00:30:19, 00:25:44) | Host's informal test |
The host's latency result sits uneasily with Google's claim that streaming runs at "sub-second" speed (Google blog). The two may have used different endpoints. No independent benchmark of the VLA itself was found.
Context and prior work
- Gemini Robotics 2 reuses "motion transfer" techniques from Gemini Robotics 1.5 (DeepMind).
- Running on flat ground is easy to train in simulation. Manipulation involving contact, such as folding cloth, is still hard to simulate (00:08:15).
- On training data, she expects a mix: teleoperation (precise but hard to scale), UMI handheld-gripper data, and egocentric human video (scalable but noisy).
Read the full section
- Gemini Robotics 2 reuses "motion transfer" techniques from Gemini Robotics 1.5 (DeepMind).
- Gopalakrishnan downplays China's viral robot races. Running on flat ground is easy to train in simulation. Manipulation involving contact, such as folding cloth, is still hard to simulate (00:08:15).
- On training data, she expects a mix: teleoperation (precise but hard to scale), UMI handheld-gripper data, and egocentric human video (scalable but noisy). This contrasts with NVIDIA's Jim Fan, who argues that egocentric video will dominate (01:15:56).
Limitations, safety and contested findings
- Benchmark success rates vary widely by task and robot body. Many are well below what deployment requires (DeepMind).
- The API docs warn that the model can hallucinate. They also make developers responsible for safety around the robot (docs).
- Gopalakrishnan separates operational safety from goal-level alignment. One example: a robot falling over is dangerous even when nothing is wrong with its goals.
Read the full section
- Benchmark success rates vary widely by task and robot body. Many are well below what deployment requires (DeepMind).
- The API docs warn that the model can hallucinate. They also make developers responsible for safety around the robot (docs).
- Gopalakrishnan separates operational safety from goal-level alignment. One example: a robot falling over is dangerous even when nothing is wrong with its goals. Another: a robot whose vision is blocked should ask for help rather than carry on (01:03:27).
- She describes the robot's gestures as natural rather than preprogrammed. They are still prompted by a model, so they are not fully emergent (01:10:10).
Business and practitioner implications
- Teams can start now by prototyping orchestration with ER 2 and their own robot APIs or VLAs declared as tools.
- Near-term deployments are most likely in pick-and-place work and in tasks where mistakes can be retried (00:14:57, 00:54:06).
- Fine-tuning through the managed API is workable without access to model weights, per the CLIFT preprint.
Read the full section
- Teams can start now by prototyping orchestration with ER 2 and their own robot APIs or VLAs declared as tools. Expect weaker results on robot APIs that are not widely used (00:43:51).
- Near-term deployments are most likely in pick-and-place work and in tasks where mistakes can be retried (00:14:57, 00:54:06).
- Fine-tuning through the managed API is workable without access to model weights, per the CLIFT preprint. The trade-off is that the action models stay locked into Google's platform (arXiv).
Sources
Read the full section
- Cognitive Revolution: One Brain, Any Body (transcript)
- Google DeepMind: Gemini Robotics 2 brings whole body intelligence to robots
- Google: Introducing Gemini Robotics ER 2
- Gemini API docs: Gemini Robotics ER
- The Robot Report: Gemini Robotics 2 enables full body control
- arXiv 2607.29172: CLIFT
- Wikipedia: Gemini Robotics
The source trail.
Sources (7)
One Brain, Any Body: Google DeepMind's Keerthana on Gemini Robotics 2, Cross-Embodiment & Humanoids
Transcript retrieved via published_transcript; language en. Timestamped text, not direct audiovisual review. Source: https://cognitiverevolution.ai/one-brain-any-body-google-deepmind-s-keerthana-on-gemini-robotics-2-cross-embodiment-humanoids. Automatic captions/transcription may contain errors.
cognitiverevolution.ai