Text to video is next
Meta's Make-A-Video generates short clips from a sentence. They're five seconds long, low resolution and physically wrong in instructive ways. The hard parts are consistency, continuity and physics.
Meta announced Make-A-Video yesterday. Type “a teddy bear painting a portrait” or “a dog wearing a superhero cape flying through the sky” and it produces a clip of a few seconds. In April I guessed the first text-to-video models would show up this year, short, low resolution and physically wrong in amusing ways. That’s a fair description of what I’m looking at, and it’s still remarkable.
The approach is clever about data. There are billions of captioned images on the internet, but far fewer captioned videos. So Make-A-Video starts with a text-to-image diffusion model that already knows what things look like and how they’re described. Then it extends the model’s layers to work across time, with temporal convolutions and attention, and trains those new parts on unlabeled video, which is plentiful, to learn how things move. Text teaches the appearance, and video without text teaches the motion. Then separate models interpolate between frames and increase resolution.
It’s the same division I’d expect people to use for a while: learn the look of the world from images and language, learn its motion from raw video.
Watching the samples closely, you can see what’s easy and what’s hard. Appearance is good. Style is good. Simple motion, like a flag waving, water rippling or a camera slowly panning, is fine. What falls apart is anything that requires the model to keep track of what’s happening. Objects melt into each other. A dog’s legs multiply and merge as it runs. A hand holding a brush doesn’t stay attached to the brush. Clips are short, a few seconds, partly because consistency degrades the longer they go.
I think the problems will be solved in roughly this order. Resolution and length first, since that’s mostly compute and engineering. Then identity consistency, so the same character looks the same throughout a clip and across clips. Then camera and scene consistency, so a room stays the same room when the camera moves. The hardest will be physics and causality: objects that have mass, collisions that make sense, hands that grip things properly. Those require the model to have learned how the world behaves, on top of how it looks.
The other thing I’d watch is control. A creator making a video wants to say more than one sentence. They want to choose the shot, the timing, the camera, the pacing, and cut between shots with continuity. Right now these models give you one uncontrollable clip. The distance between “generate a clip” and “make a film” is huge, and I think it’ll be closed by tools that combine generation with editing, not by a single model that does everything from one prompt.