Mobile ALOHA cooks shrimp for $32,000
Stanford's two-armed robot on a wheeled base learned to cook, wipe spills and call an elevator from about fifty demonstrations per task. Cheap hardware plus teleoperation is changing robot data.
A Stanford team posted Mobile ALOHA last week, and the videos went everywhere. A two-armed robot on a wheeled base sautés shrimp and flips it, wipes up spilled wine, calls and enters an elevator, pushes in chairs, and rinses a pan. The whole system costs about $32,000, including the laptop that runs it.
A lot of the viral clips were teleoperated, meaning a person was driving the robot, which the authors say clearly and which a lot of reposts leave out. The autonomous results are more modest but still impressive: after about fifty demonstrations per task, the robot does these tasks on its own with reasonable success rates.
I think the most important part is the data recipe, not the cooking.
The teleoperation rig is the clever piece. A person stands behind the robot, physically attached to its base, and holds two small “leader” arms that the robot’s “follower” arms mirror. So the operator walks the robot around and moves its arms naturally, with both hands, and every demonstration records video from the cameras and the joint positions of both arms and the base. Collecting fifty demonstrations of a task takes an afternoon, not a month.
They also found that co-training helps a lot. They mixed their new mobile demonstrations with an existing dataset of demonstrations from the static ALOHA setup, which covers different tasks. The policy learned some general manipulation skills from the static data that transferred to the new tasks, and success rates went up substantially, up to 90% on some tasks.
Put together, this says something about where robot learning is heading. The bottleneck has been data. Language models learn from the internet, but no internet of robot actions exists. Mobile ALOHA’s answer is to make the collection cheap: affordable hardware, an intuitive teleoperation rig that anyone can use, and a policy that improves by pooling data across tasks and robots.
I’m less sure it’s the whole answer. Fifty demonstrations per task is cheap per task, but the world has millions of tasks. Teleoperation also only captures what a person does through the robot’s body. It’s shaped by the rig, and it doesn’t include touch or force. It’s a big step toward useful household robots for specific jobs.
For general physical intelligence, I still think we’ll need the other half: robots that practice on their own, mostly in simulation, learning the physics of their bodies and objects the way a child does, and use a small amount of human demonstration to learn what we want from them. The demonstrations would then only need to cover what’s specific to each task, which is a much smaller amount of data.