Vinson·Li

Essay No. 112

Optimus sorts blocks. A toddler falls down a thousand times

Tesla's new video shows Optimus sorting colored blocks with a neural network trained end to end. Curated robot demos and messy child learning, and which one scales.


Tesla posted a new Optimus video last Saturday. The robot sorts blue and green blocks into trays, corrects itself when a person moves a block, and then does a few yoga poses, balancing on one leg. Tesla says the sorting is done by a neural network trained end to end, from camera video to joint commands, running on the robot.

It’s a real improvement over last year’s slow walk. The hand movements are smooth, and the correction when someone knocks a block over looks like a system reacting, not a script. The yoga shows good whole-body balance control.

I watched it twice, and then went to a friend’s house for dinner, where their one-year-old was learning to stand. And I kept comparing the two.

The robot’s learning looks like this: a task is chosen, a lot of demonstrations are collected, probably by people teleoperating the robot, and a network is trained to imitate them. The result is good at that task, on that table, with those blocks. Each new task needs more demonstrations. The video is, understandably, curated. We see the successes.

The toddler’s learning looks like this: no task, no demonstrations, no curation. She pulls herself up on the coffee table, lets go, wobbles, and falls on her bottom. Again. And again. Nobody is labeling anything. She’s learning balance, the strength of her own legs, the fact that the table holds her weight and the cushion doesn’t, what happens to a cup when you knock it off the table (her parents know this one well). Every fall teaches her about gravity, her body and the world at once. And in a few months she’ll use that same body knowledge to walk, climb, carry things and throw.

I know it’s an unfair comparison. The toddler has had millions of years of evolution tuning her brain and body for exactly this, and several hundred million years of vertebrate motor control built in. Robots have a few years of engineering. But the difference in how they learn seems important to me. One learns specific skills from curated data. The other learns a general model of her body and the world from undirected play, with curiosity as the only teacher, and then gets specific skills almost for free.

I’ve written for years that the second way is the one that scales, and I still think so. The current generation of robot learning, imitating demonstrations task by task, will produce useful robots for narrow jobs. The generation after that will need robots that learn the way she does: a body with realistic limits, lots of senses including touch, a very large number of safe falls in simulation, and the motivation to try things just to see what happens.

For now, the robot sorts blocks better than the toddler does. I’d bet on the toddler.

Fin.

Add a comment

Comments

Plain text

  • Loading comments…