Learning like a baby: a plan for an embodied world model
Google DeepMind's Gemini Robotics ER 2 gives robots a better high-level brain. The part I still think nobody has built is the body-first learning underneath. Here's the research plan I'd run.
Google DeepMind released Gemini Robotics ER 2 yesterday, an embodied reasoning model that acts as the high-level brain for a robot: it understands the scene from continuous video, plans multi-step tasks, tracks progress, coordinates several robots, and hands motor execution to lower-level action models. It’s available to developers through the Gemini API. (I work at Google, not on this, and I’m going by the public announcement.)
It’s a good example of where the field is: very strong understanding and planning on top, and learned motor policies below, trained mostly from demonstrations. That’s the Helix split I wrote about last year.
What I think is still missing sits underneath both: a model of the body and the physical world learned by the agent itself, through its own experience, the way children learn it. I’ve been circling this on this blog for over a decade, from DQN failing Montezuma’s Revenge in 2015 to curiosity in 2017, Dactyl’s hand in 2018, and “A billion dollars for JEPA” in March. So here’s the concrete research plan I’d run, and have started prototyping on rented GPUs in my own time.
The body. A simulated humanoid built on real robot constraints: link masses, joint ranges, torque and speed limits modeled on current humanoid designs like Optimus, not an idealized ragdoll. Two articulated hands with a realistic number of degrees of freedom. No grippers. Hands are where most of what we understand about objects comes from.
The senses. Vision from head cameras, but also proprioception (joint angles and velocities), force and torque at the joints, and touch across the fingertips, palms and body surface, since contact is where physics shows up. Sound, because impacts and scraping carry information about materials. All of it goes into one sequence of tokens, the way Gato and Perceiver showed in 2021 and 2022.
The curriculum. Babies don’t start by running. They start by discovering they have a body. First, lying in a crib-like space, learn how your own limbs respond to your motor commands. Then rolling, then crawling, pulling up, standing, falling, walking, then jumping and running. Each stage is enabled by what the model learned in the previous one. Hands come in early: reaching, grasping, dropping, banging objects together.
The reward. Mostly intrinsic. The agent is rewarded for learnable surprise, meaning its world model’s prediction error on things it can influence, measured in representation space so it doesn’t get hypnotized by noise. That’s the noisy TV lesson from 2018. A small number of extrinsic goals are introduced only once the body is competent.
The world model. A JEPA-style predictor over all senses: given the current representation and an action, predict the next representation, across vision, touch, force and sound. Planning happens inside it, MuZero-style. Video pretraining, like V-JEPA 2’s, provides a prior for what the world looks like. The body provides physics.
The playground. Procedurally generated environments with objects that vary in mass, friction, softness and shape, and consistent rules, like Breath of the Wild’s “chemistry engine” that I wrote about in 2017. Eventually, generated worlds from the kinds of models World Labs and DeepMind are building, with physics attached.
The question I can’t answer yet is scale: how big a model and how much simulated experience would it take before something like intuitive physics, and then general manipulation, emerges? A child needs about a year of waking experience to walk well. Simulation can run far faster than real time, in parallel. My guess is the compute is within reach. The hard part is simulating contact and touch faithfully, which is where physics engines are weakest.
I’ll write up results as they come, including the failures.