Genie 2 and World Labs in the same week
DeepMind's Genie 2 turns one image into a playable 3D world, and Fei-Fei Li's World Labs turns one image into a 3D scene you can walk through. Worlds are the next medium after text, images and video.
Two announcements, two days apart, about the same idea from different directions.
On Monday, World Labs, the startup Fei-Fei Li co-founded this year with a lot of funding, showed its first results: give it a single image, a photo or a painting, and it generates a 3D scene you can move around in, in your browser. The demos include walking into Van Gogh’s Café Terrace at Night and looking around the corner. The generated scenes stay consistent as you move, since the model produces actual 3D geometry that’s rendered in real time. You can apply effects like depth of field and change the camera freely.
On Wednesday, Google DeepMind showed Genie 2. Give it a single image, and it turns it into a playable 3D environment that you, or an AI agent, can control with a keyboard and mouse for up to about a minute. It models physics like gravity, water and smoke, lighting and reflections, object interactions, character animation, and it remembers parts of the world that go out of view and come back. It’s a video model conditioned on actions, the approach of the original Genie in March and GameNGen in August, now in 3D and on arbitrary images.
The two take different approaches, and I think the difference is instructive.
World Labs produces an explicit 3D scene. There’s persistent geometry that exists independently of where you look. That gives you strong consistency: the room is the same room when you turn around, because it’s actually a room. It also makes it easy to use in normal graphics tools. The limits right now are that the scenes are fairly small and mostly static, and the interaction is mostly camera movement.
Genie 2 generates the next frame from the previous frames and your action. There’s no explicit 3D. The world only exists as the model’s prediction. That gives you interaction and dynamics: things move, fall and respond to you. But consistency has to be learned and maintained by the model, and it drifts after a minute or so.
I’ve written for years about representation: meshes versus implicit fields versus Gaussian splats, pixels versus latent predictions. This is the same question at a higher level. Should a world model hold an explicit representation of the world, or should the world only exist as predictions? My guess is the answer ends up a hybrid: explicit structure for the things that must persist, like geometry and identity, and learned dynamics for the things that change, like physics and behavior.
Either way, what both demos made clear to me is that after text, images, audio and video, the next generative medium is worlds, meaning spaces you can enter and act in, not frames you watch. That matters for games and film. It matters even more for training embodied agents, which need somewhere to practice. DeepMind says as much in the Genie 2 post, and it’s the application I’m most interested in.