Vinson·Li

Essay No. 105

Control beats prompts

ControlNet lets you steer Stable Diffusion with a pose skeleton, a depth map or an edge sketch. Creators want to set the structure directly, and this gives them a way to.


Ten days ago Lvmin Zhang and Maneesh Agrawala at Stanford posted “Adding Conditional Control to Text-to-Image Diffusion Models,” and the release that came with it, ControlNet, is already all over the Stable Diffusion community. You give it a text prompt plus a structural input, like a human pose skeleton, a depth map, an edge map, a rough scribble or a segmentation map, and it generates an image that follows both. Draw a stick figure in a pose and get a photo of a person in exactly that pose. Take the edges of your living room and get it re-imagined in any style, with the furniture where it actually is.

The technique is neat. They take the pretrained Stable Diffusion network and freeze it. They make a trainable copy of its encoder blocks, which receives the structural input. The copy’s outputs are added back into the frozen network through “zero convolutions,” one-by-one convolutions whose weights start at exactly zero. At the beginning of training, the copy contributes nothing, so the model behaves exactly like the original and nothing gets damaged. As training goes on, the zero convolutions learn to let the control signal in. It can be trained on a single consumer GPU for some conditions, with a modest dataset.

Why I think this matters more than the next improvement in prompt following: text is a bad language for structure. You can write “a woman sitting on a bench, facing left, with her right hand raised” and the model will get some of it right. A pose skeleton gets all of it right in one gesture. Creators, whether illustrators, designers or filmmakers, think in composition, layout, pose and camera. They want to specify those precisely and let the model fill in the rest. A prompt-only tool forces them to describe what they could show, and then gamble.

This is also the key to video. The biggest problem with generated video, which I wrote about in September, is consistency across frames: characters that change, objects that drift, motion that makes no sense. If you can condition each frame on structure, like a pose sequence from a real performance, a depth map from a 3D scene, or the edges of an existing video, much of that consistency comes from the conditioning instead of having to be invented by the model. People are already doing crude versions of this, running ControlNet frame by frame over a video, with a lot of flicker, but the direction is obvious.

My prediction is that the successful creative AI tools won’t be “type a sentence, get a result.” They’ll combine generation with the kind of control creators already use: sketches, reference poses, camera paths, timelines and existing footage. The model supplies the rendering, and the person supplies the structure and the taste.

Fin.

Add a comment

Comments

Plain text

  • Loading comments…