Vinson·Li

Essay No. 133

62 hours of robot data

Meta's V-JEPA 2 learns a world model from a million hours of video, then learns to plan robot actions from 62 hours of robot data. Promising, and still missing touch.


Meta released V-JEPA 2 last week. It’s the most complete version yet of the approach I’ve been following since LeCun’s position paper in 2022 and I-JEPA in 2023: learn to predict in representation space, from video, without labels, then use that for understanding and for acting.

It’s trained in two stages.

The first stage is observation. A large video encoder, about a billion parameters, is trained on over a million hours of internet video plus images, with the JEPA objective: mask out parts of a video in space and time, and predict the representations of the missing parts from the visible ones. No pixel reconstruction and no captions. The result is a representation of video that, according to Meta’s evaluations, is strong at understanding motion and at anticipating what happens next in a scene.

The second stage is action. They freeze that encoder and train an action-conditioned predictor, V-JEPA 2-AC, on 62 hours of robot videos from the public DROID dataset, which records a robot arm’s camera views together with its actions. The predictor learns: given the current representation and an action, what will the representation be after the action?

Then they put it on real Franka robot arms in two labs that weren’t in the training data, and give it a goal image, like the object placed at a target spot. The robot plans by imagining: sample candidate action sequences, predict where each one leads in representation space, pick the one whose predicted end state is closest to the goal’s representation, execute the first step, and replan. Meta reports success rates between 65% and 80% on pick-and-place with objects the model hasn’t seen, with no task-specific training and no reward.

This is a world model in the full sense: it predicts consequences of actions and supports planning. And it shows the division of labor I’ve expected for a while. Most of the knowledge about how the world looks and moves can come from passive video, which is abundant. A small amount of interaction data is enough to connect that knowledge to a specific body’s actions. 62 hours is tiny compared to the million hours of watching.

What’s missing is the part I keep coming back to on this blog. The robot has a camera and a gripper. It has no touch, no force sensing in the representation, and no dexterous hand. Pick-and-place with a parallel gripper is the easiest kind of manipulation, and the representation is built entirely from what things look like. Picking up a soft object, knowing how hard to squeeze a cup, feeling when a screw starts to catch: none of that is in video, and none of it can be learned by watching.

My guess for the next step: the same recipe, with video-pretrained representations extended to include touch, force and proprioception from robots with real hands, trained with much more interaction, a lot of it in simulation. The watching gives you the world’s appearance and rough dynamics. The body gives you the rest.

Fin.

Add a comment

Comments

Plain text

  • Loading comments…