Gemini Robotics: three models, and a published record with almost no numbers in it
What Google DeepMind's Gemini Robotics pages actually state about its action and reasoning models, and how little of it carries a measurement.
Google DeepMind describes Gemini Robotics as two kinds of model with one job. A vision-language-action (VLA) model converts vision and language into motor control, an embodied reasoning (ER) model plans and coordinates, and the model page says they "work together to interact with the physical world. Each has a specialist role, but they operate as one powerful and versatile system." The current generation is called Gemini Robotics 2, and the page's headline pitch is "The intelligence layer to power any kind of robot."
The striking thing about the published material is what is absent from it. The Gemini Robotics model page carries no release date, no version history, no model card, no specification block and no evaluation figure of any kind. The one dated page in the family, the March 2025 launch post, makes two performance claims and cites its technical report for both without printing a number on the page. This post sets out what the pages state, and marks clearly where a statement is the lab describing its own product.
Companion reading: Foundation models and VLAs in robotics, imitation learning, edge AI and robot compute, humanoid robot hardware.
Table of contents
- Key takeaways
- The stack, and how it got here
- What the record actually states
- What you can use today
- Demonstrated versus claimed
- What is missing from the record
- What to watch
- Frequently asked questions
- Changelog
The stack, and how it got here
The line starts on 12 March 2025 with two models built on Gemini 2.0. The first is Gemini Robotics, a VLA "built on Gemini 2.0 with the addition of physical actions as a new output modality for the purpose of directly controlling robots." The second is Gemini Robotics-ER, described as "a Gemini model with advanced spatial understanding, enabling roboticists to run their own programs using Gemini's embodied reasoning (ER) abilities."
The launch post pairs those with a hardware partnership and a tester programme. DeepMind says it is "partnering with Apptronik to build the next generation of humanoid robots with Gemini 2.0," and separately that it is "working with a selected number of trusted testers to guide the future of Gemini Robotics-ER." The trusted tester access described there is for the ER model. Later on the same page it names those testers as Agile Robots, Agility Robots, Boston Dynamics and Enchanted Tools.
Two successors appear on that page only as related-post entries, giving a title and a month. "Gemini Robotics On-Device brings AI to local robotic devices" is dated June 2025, and "Gemini Robotics 1.5 brings AI agents into the physical world" is dated September 2025. Neither entry carries any technical detail, and no further release is listed.
The current generation is documented on the model page, which names three models and gives each a one-line description. Gemini Robotics 2 is "Our most advanced vision-language-action model (VLA) that converts vision and language input into motor control, enabling a robot to take action." Gemini Robotics ER 2 is "Our embodied reasoning model: capable of reasoning within physical spaces to make detailed plans, coordinating with humans and other robots." Gemini Robotics On-Device 2 is "A lightweight version of our VLA model, optimized to run locally on robotic hardware." The page gives no date for any of them, and no base model.
What the record actually states
No success rate, benchmark score, accuracy, latency figure, parameter count or context length appears anywhere on either page. There is no evaluation chart, no comparison table and no named competitor model. Video demonstrations are embedded, and none of them is accompanied by a reported result.
That leaves two performance claims, both from March 2025, both about the first generation, and both deferred to a document off the page. DeepMind writes that "in our tech report, we show that on average, Gemini Robotics more than doubles performance on a comprehensive generalization benchmark compared to other state-of-the-art vision-language-action models." Separately, on the ER model running perception through to code generation, it says that "in such an end-to-end setting the model achieves a 2x-3x success rate compared to Gemini 2.0."
Both are the lab reporting on its own model. The tech report link resolves to arXiv 2503.20020. As printed on the blog page, neither claim states the benchmark, the task count or the absolute rate, so neither can be checked without opening the report. The benchmark named in the first claim is not identified.
Nothing comparable exists for the current generation. There is no equivalent claim, sourced or otherwise, for Gemini Robotics 2, ER 2 or On-Device 2 on the model page.
What you can use today
The model page offers one live route. A "Try Gemini Robotics ER 2" button leads to Google AI Studio, and the model identifier gemini-robotics-er-2-preview sits inside that link's URL. The page text never describes an availability tier, a preview status, pricing or terms, so the identifier in the link is the whole of the published detail.
Everything else runs through a single funnel. Under "Become a trusted tester" the page says it is working with "100+ trusted testers" to "deploy, test, and guide the future of Gemini Robotics in the real world," followed by "Join the waitlist for early access." That waitlist covers Gemini Robotics as a whole. The page does not state separate access terms for the VLA or the on-device model.
Neither page offers weights, a download, an SDK or a repository for any model in the family. The open artefacts named are from March 2025: the technical report at arXiv 2503.20020, and the ASIMOV dataset, which the launch post introduces as "a new dataset to evaluate and improve semantic safety in embodied AI and robotics."
Named partners are few. The model page lists Agile Robots, Apptronik and Boston Dynamics under "Research partners" and names none of the 100-plus testers. The hardware named in the material is all from March 2025: DeepMind says it trained "primarily on data from the bi-arm robotic platform, ALOHA 2," demonstrated control of "a bi-arm platform, based on the Franka arms used in many academic labs," and specialised the model for "the humanoid Apollo robot developed by Apptronik." No hardware is named for the current generation.
Demonstrated versus claimed
Demonstrated, in the sense of a figure a reader can check against the page: nothing. Neither page publishes one.
Claimed, in the lab's own words and with no evaluation attached, from the model page: "The intelligence layer to power any kind of robot"; "Our most advanced vision-language-action model"; "Can be adapted to any bi-arm robot in just a few hours"; "Controls complex humanoid hands and parallel grippers to unlock a new level of physical dexterity"; "Controls entire humanoid bodies from feet to fingertips"; and, on multi-robot work, "Enables different types of robots to communicate and work together to solve complex workflows a single robot could not do alone." These are product descriptions written by the vendor. Read them as positioning until an evaluation appears.
A middle category holds the two March 2025 claims above. They are quantified, they are self-reported, and they point to a technical report the pages do not summarise.
The launch post is candid about one thing at least: it frames embodied reasoning as a hard problem and describes classic safety measures, collision avoidance, contact force limits and dynamic stability, as the responsibility of "'low-level' safety-critical controllers, specific to each particular embodiment" that the ER model interfaces with. The motor safety layer is explicitly somebody else's.
What is missing from the record
Everything in this section is an absence from the two pages fetched for this post, checked against them directly.
Neither page gives a release date for the current generation, a version history, or a base model for any of the three current models. There are no parameter counts, no context window or token limits, no input and output modality specification, no latency, no control frequency and no pricing or rate limits.
The on-device story is the thinnest. The phrase "optimized to run locally on robotic hardware" appears with no compute target, no accelerator, no model size and no memory footprint. For a model whose selling point is local execution, that is a large blank.
Training data is undescribed for the current generation. Nothing states teleoperation hours, episode counts, robot counts or scene counts, and no training infrastructure is named.
On evaluation, there is no benchmark table, no competitor baseline, no trial count, no error bars and no model card for any of the three. No named production customer, deployment count, safety incident record, reliability figure or compliance certification appears in either page.
What to watch
Whether a technical report or model card lands for the current generation with the figures the pages omit, in the way the March 2025 release at least pointed to a report.
Whether the model page ever dates its releases. A model family with three named generations and no date on the page is hard to track and easy to misreport.
Whether the VLA gets a route beyond the waitlist. Today the reasoning model has a button and the action models have a form.
Whether the dexterity and whole-body bullets acquire an evaluation. "A new level of physical dexterity" is a claim that a success rate on a named task would settle.
Whether the "100+ trusted testers" figure converts into named deployments.
Frequently asked questions
Can I use Gemini Robotics today?
The model page links Gemini Robotics ER 2 into Google AI Studio, and the model identifier in that link is gemini-robotics-er-2-preview. For the action models the page offers a trusted tester waitlist and nothing else.
Are the weights open? Neither page offers weights or a download for any model in the family. The open artefacts named are the March 2025 technical report at arXiv 2503.20020 and the ASIMOV safety dataset.
What is the difference between the VLA and ER models? DeepMind describes the VLA as converting vision and language input into motor control so a robot can take action, and the ER model as reasoning within physical spaces to make detailed plans and coordinate with humans and other robots.
What are the current models built on? The pages do not say. The first generation, from March 2025, is stated to be built on Gemini 2.0. No base model is given for Gemini Robotics 2, ER 2 or On-Device 2.
How well does it work? On the two pages reviewed, there is no published figure to answer that. The only quantified claims are DeepMind's own from March 2025, about the first generation, and both defer to the technical report.
Which robots has it been shown on? The March 2025 post names ALOHA 2 as the primary training platform, a bi-arm platform based on Franka arms, and Apptronik's Apollo humanoid. No hardware is named for the current generation.
Who are the partners? The model page lists Agile Robots, Apptronik and Boston Dynamics as research partners. The March 2025 post named Agile Robots, Agility Robots, Boston Dynamics and Enchanted Tools as trusted testers for the ER model.
Related guides
- NVIDIA Isaac GR00T: The Open Humanoid Model That Ships With Its Own Compute Bill
- Tesla Optimus Review: The Scorecard, Not the Demo Reel
- Figure 03 Review: The Only Humanoid With a Repeat Customer
- Wayve: the end-to-end driving bet, and what it has actually shown
- Waabi: a simulation-first bet on autonomous trucking
- Foundation Models & VLAs for Robotics: The Ultimate Guide