Vinson·Li

Essay No. 142

A billion dollars for JEPA

Yann LeCun's AMI Labs raised $1.03 billion to build world models on the JEPA framework. Revisiting what I wrote in 2022, and the one thing I still think the plan is missing.


AMI Labs, the company Yann LeCun started after leaving Meta, announced on Monday that it raised $1.03 billion in a seed round at a $3.5 billion pre-money valuation. It’s based in Paris, with Alexandre LeBrun as CEO and LeCun as executive chairman, and it’s building world models on JEPA, the Joint Embedding Predictive Architecture. It’s reportedly Europe’s largest seed round ever.

In July 2022 I wrote a post called “LeCun’s path, read carefully,” about his position paper. I said the core idea, predicting in representation space instead of generating every detail, seemed right, and that I’d push back on how much can be learned from watching, because the most important things about physics come from acting and feeling. Almost four years later, I’d keep both halves of that.

The first half has held up well. I-JEPA in 2023 showed that predicting hidden regions in latent space learns strong image representations efficiently. V-JEPA 2 last June showed the approach on video at scale, and then planning robot actions from 62 hours of robot data. Meanwhile the frame-generating world models, from Sora to Genie 3, keep showing both how much can be learned from pixels and how much effort goes into rendering detail that doesn’t matter for decisions. JEPA’s bet, that a planning model should predict what matters and ignore the rest, looks more reasonable to me every year.

The second half I feel more strongly about now. A world model for physical intelligence has to learn what a body can do: how heavy things are, how hard to grip, how to stay balanced, what happens at contact. Those aren’t in video. They’re in touch, force, proprioception and the consequences of your own movements. V-JEPA 2’s robot arm had a camera and a gripper. Every world model I’ve written about in the last two years is, in the end, a camera moving through a space, or a gripper grabbing in a space. None of them has a body with the constraints of a real one.

If I had a small fraction of that billion dollars, this is what I’d spend it on. A simulated humanoid with realistic mass, joint limits, torque limits and speeds, based on real robot designs, with articulated hands, not grippers, and with touch and force sensing across the skin as well as cameras. Let it learn its own body first, the way a baby does, from rolling to crawling, standing, walking and eventually running and jumping, with curiosity as the main reward, as in the 2017 and 2018 papers I wrote about. Then use a JEPA-style predictor across all of those senses as the world model it plans in. Video pretraining supplies the look of the world. The body supplies the physics.

I’m glad the bet is being made at this scale. I think the teams that treat the body as part of the world model, instead of an output device attached to it later, are the ones who’ll get furthest.

Fin.

Add a comment

Comments

Plain text

  • Loading comments…