Genie 3 remembers where you painted the wall
DeepMind's Genie 3 generates interactive worlds in real time at 720p that stay consistent for minutes. Persistence is the test for a world model, and it just got much better.
On Tuesday Google DeepMind announced Genie 3. It generates interactive environments from a text prompt, in real time, at 720p and 24 frames per second, and you can move around in them for a few minutes. I work at Google, but wasn’t involved in Genie 3. This post is based on the public announcement and videos.
I’ve followed this line from the first Genie, a year and a half ago, which made low-resolution 2D platformers at one frame per second, to GameNGen’s Doom and Genie 2’s minute-long 3D worlds in December. Genie 3 is a big jump on every axis. But the demo that tells you the most is the painting one.
Someone is in a room, and they paint on a wall with a paint roller. They turn away, look at other parts of the room, and then turn back. The paint is still there, exactly where they put it.
That sounds trivial, and it’s one of the hardest things for a model like this. Genie 3 has no explicit 3D representation of the room. Every frame is generated from the previous frames and your actions. When you turn away from the wall, it’s no longer on screen, so the only place the paint exists is in the model’s memory of what it generated before. For earlier frame-by-frame models, anything off screen tended to drift, change or disappear. DeepMind says Genie 3’s visual memory reaches back about a minute, and that consistency holds over several minutes of interaction.
In my December post comparing Genie 2 with World Labs, I said persistence is where frame-generating models are weakest compared to explicit 3D scenes, and that the answer would probably be a hybrid. Genie 3 closes a lot of that gap without any explicit 3D at all. That surprised me. It suggests that with enough scale, consistency can emerge from learned memory.
The other new feature is “promptable world events.” While you’re exploring, you can type something like “a herd of deer runs across the path” or “it starts to snow,” and the world changes accordingly. That’s a big deal for the use DeepMind emphasizes: training agents. If you want an agent to learn to handle surprises, you need a world where you can make surprising things happen on demand.
The limits are clear. The action space is small, mostly moving and looking. Multiple agents interacting is hard. Text in the world is unreliable. And a few minutes is still short. It’s not released publicly, only to a small group of researchers.
But the trajectory over twenty months is what matters: from 2D at one frame per second, to a real-time, persistent 3D world you can change with a sentence. If it continues, models like this become the simulators embodied agents train in, and possibly a new medium for people too. What’s still missing for the agent case is a body with realistic physical limits inside the world. Right now the agent is a camera floating through it.