On world-models
- Learning like a baby: a plan for an embodied world model
Google DeepMind's Gemini Robotics ER 2 gives robots a better high-level brain. The part I still think nobody has built is the body-first learning underneath. Here's the research plan I'd run.
3 min reads likes comments - A billion dollars for JEPA
Yann LeCun's AMI Labs raised $1.03 billion to build world models on the JEPA framework. Revisiting what I wrote in 2022, and the one thing I still think the plan is missing.
2 min reads likes comments - A weekend with the World API
World Labs opened an API for Marble. I spent the weekend building a small workbench for rendering its worlds as Gaussian splats and authoring camera paths through them. What broke: scale, coordinates and drift.
3 min reads likes comments - LeCun leaves, Marble ships
In one week World Labs launched Marble, Yann LeCun confirmed he's leaving Meta to start a world-model company, and Google shipped Gemini 3. The world-model bet went public from three directions.
2 min reads likes comments - Genie 3 remembers where you painted the wall
DeepMind's Genie 3 generates interactive worlds in real time at 720p that stay consistent for minutes. Persistence is the test for a world model, and it just got much better.
2 min reads likes comments - 62 hours of robot data
Meta's V-JEPA 2 learns a world model from a million hours of video, then learns to plan robot actions from 62 hours of robot data. Promising, and still missing touch.
2 min reads likes comments - Genie 2 and World Labs in the same week
DeepMind's Genie 2 turns one image into a playable 3D world, and Fei-Fei Li's World Labs turns one image into a 3D scene you can walk through. Worlds are the next medium after text, images and video.
2 min reads likes comments - A neural net runs Doom
Google Research's GameNGen simulates Doom in real time with a diffusion model: no game engine, just next-frame prediction conditioned on your inputs. Interactive video is a world model.
2 min reads likes comments - Genie learned to play from videos with no controls
DeepMind's Genie learned a controllable world model from 2D platformer videos with no action labels, by inferring eight latent actions on its own. The model learns the controls as well as the game.
2 min reads likes comments - "World simulator" is doing a lot of work
OpenAI's Sora generates minute-long videos that look astonishing. Its technical report calls it a world simulator. The samples show both why that's tempting and why it isn't true yet.
2 min reads likes comments - Predicting in representation space
Meta's I-JEPA is the first concrete result from LeCun's world model agenda. It learns image representations by predicting hidden regions in latent space, with no augmentations and no pixel reconstruction.
2 min reads likes comments - LeCun's path, read carefully
Yann LeCun's position paper argues that intelligence needs world models that predict in representation space, not pixels. Where I agree, and where I'd push back: bodies and hands.
2 min reads likes comments - MuZero learns the rules it isn't given
DeepMind's MuZero plays Go, chess, shogi and Atari at top level without being told the rules. It plans inside a model it learned, and the model only predicts what matters.
2 min reads likes comments - An agent that dreams its own racetrack
Ha and Schmidhuber's World Models compresses what an agent sees, learns to predict what happens next, and trains a tiny controller inside its own dream. The most important paper I've read this year.
2 min reads likes comments