Genie learned to play from videos with no controls
DeepMind's Genie learned a controllable world model from 2D platformer videos with no action labels, by inferring eight latent actions on its own. The model learns the controls as well as the game.
A week after Sora, DeepMind published Genie, “Generative Interactive Environments.” It gets less attention than Sora because the output is low-resolution 2D platformer games. I think it’s more interesting from a world-model point of view.
Genie was trained on about 200,000 hours of gameplay videos of 2D platformers from the internet, filtered down to around 30,000 hours of usable footage. There are no action labels. Nobody recorded which buttons were pressed, the same problem OpenAI’s VPT had with Minecraft in 2022. VPT solved it with a small labeled dataset and an inverse dynamics model. Genie solves it without any labels at all.
It has three parts.
A video tokenizer compresses frames into discrete tokens.
A latent action model looks at a frame and the next frame and asks: what action could explain the change? It’s forced to answer with one of only eight discrete codes. Because it’s so constrained, and because it’s trained so that the next frame can be predicted from the current frames plus the code, it learns codes that correspond to consistent, controllable changes. In practice, codes end up meaning things like “move left,” “move right,” “jump,” across completely different games. Nobody told it that platformer characters have a small set of consistent controls. It discovered that the most useful way to explain frame-to-frame change with eight options is to find the player’s controls.
A dynamics model predicts the next frame’s tokens from past frames and a latent action.
At inference, you give it one image, even a hand-drawn sketch or a photo it has never seen, and it turns it into a playable environment. You pick one of the eight actions each step, and it generates the next frame. Press “right” and your sketched character walks right, with parallax in the background.
What I like about the demo is that you can actually play it: it predicts what happens next as a function of what you do. And it learned the space of actions from passive video, with no instrumentation.
The authors also show a version trained on robot videos, with no actions, that learns latent actions resembling the robot’s movements. That’s the part that could matter a lot. If a model can learn an action space and a controllable dynamics model from unlabeled video, the internet’s video of people and machines doing things becomes training data for interactive world models, and then possibly for agents trained inside them, like the “World Models” paper did in a car game in 2018.
The limits right now: 1 frame per second of generation, low resolution, 16 frames of memory, and eight actions is a tiny space compared to a human body. But a year ago I didn’t expect an unsupervised model to discover a controller from video at all.