"World simulator" is doing a lot of work
OpenAI's Sora generates minute-long videos that look astonishing. Its technical report calls it a world simulator. The samples show both why that's tempting and why it isn't true yet.
OpenAI showed Sora on Thursday. It generates videos up to a minute long from text, at high resolution, with multiple characters, camera motion and detailed backgrounds. The woman walking down a neon-lit Tokyo street, the woolly mammoths in the snow, the drone shot of the Big Sur coastline. I’ve been following video generation closely since Make-A-Video in 2022, and this is the biggest single jump I’ve seen.
The technical report is titled “Video generation models as world simulators.” That framing deserves a close look.
What’s described, at a high level: videos are compressed by a network into a lower-dimensional latent space, over time and space. That latent video is cut into “spacetime patches,” small chunks of space over a few frames, which act as tokens. A diffusion model with a Transformer backbone (a diffusion transformer) is trained to denoise those patches, conditioned on text. It’s trained on videos at their native resolution, aspect ratio and duration, instead of cropping everything to squares. And like DALL·E 3, it uses highly detailed synthetic captions for training and a language model to expand user prompts.
The report argues that as you scale this up, 3D consistency, object permanence and some physical interactions emerge, without anything explicitly 3D or physical in the model. The camera moves and the scene stays geometrically consistent. A painter’s brushstrokes stay on the canvas. Someone takes a bite of a burger and it has a bite mark afterwards.
Then there’s the failure section, which is the most honest part and which I’d recommend anyone read. A glass falls and doesn’t shatter properly. Liquid does strange things. Someone walks on a treadmill backwards. Objects appear spontaneously, like extra wolf pups multiplying. A chair floats and deforms as if made of cloth.
Here’s how I read those failures. They’re exactly what you’d expect from a model of what videos look like, and not what you’d expect from a model of how the world works. A physics simulator can be inaccurate, but it doesn’t spawn extra mass or let a solid chair behave like cloth, because it has state: objects that persist, with properties, obeying rules. Sora has an extraordinarily good statistical model of pixels over time. When the statistics of typical videos are enough, the result looks like physics. When the situation is unusual, like a treadmill used backwards or a glass shattering in a specific way, the statistics fail and you see what’s underneath.
I’m not dismissing it. Learning this much 3D consistency and object persistence from video alone is remarkable, and it’s evidence that a lot of world knowledge is learnable from watching. But a world simulator needs to be right about counterfactuals: what happens if I push this, what happens if I don’t. A generator that produces plausible video doesn’t need to answer that, and I don’t think it does yet.
It’s also hard to steer. One prompt in, one clip out. As I wrote in November, making something with continuity across shots needs a persistent representation of the scene. Sora’s quality raises the ceiling a lot. The gap between clip and film is still there.