Vinson·Li

Essay No. 92

Diffusion is going to eat GANs

Two papers this week, GLIDE and latent diffusion, make text-to-image generation with diffusion models look practical. What denoising actually learns, and why it beats the adversarial game.


Two papers came out a week ago on the same day that I think mark a turning point. OpenAI’s GLIDE generates images from text with a diffusion model, and people rated its samples above DALL·E’s. And a group from Heidelberg and Munich posted “High-Resolution Image Synthesis with Latent Diffusion Models,” which makes diffusion much cheaper by running it in a compressed space.

I wrote about GANs in 2014 when they produced blurry digits, and about StyleGAN in 2018 when they produced photographic faces. For most of the last seven years, if you wanted high quality images from a neural network, you used a GAN. I now think that’s ending.

A diffusion model learns to reverse a process that destroys images. Take a real image and add a little Gaussian noise. Add a bit more. Keep going for a thousand steps until it’s pure static. That forward process has no learning in it. The model is trained on the reverse: given a noisy image and how much noise was added, predict the noise, so you can subtract it and get a slightly cleaner image. To generate, you start from pure static and denoise step by step, and an image emerges.

Why this beats the GAN game, as far as I can tell:

Training is stable. A GAN is two networks fighting, and a lot of GAN research was about keeping the fight from collapsing. A diffusion model is trained with a simple regression loss on one network. It just works more often.

It covers the whole distribution. GANs tend to find a few outputs that fool the discriminator and stick to them, which is mode collapse. Diffusion models are trained on every noise level of every training image, so they learn to produce the full variety of the data.

It’s easy to steer. GLIDE uses a trick called classifier-free guidance: train the model both with and without the text prompt, and at sampling time push the prediction in the direction the prompt adds. You can turn a knob for how closely to follow the prompt.

The drawback has been cost. Hundreds of denoising steps on full-resolution images is slow and expensive. That’s what the latent diffusion paper addresses. It first trains an autoencoder that compresses images into a much smaller latent representation, then runs diffusion there. Much less compute, similar quality. The authors released code and models.

My prediction: 2022 will be the year of image generation. Text-to-image models built on diffusion will get good enough that people outside research start using them daily, and at least one strong model will be released openly so anyone can run it. After that, the same approach will move to audio and video, which are the same problem with more dimensions.

Fin.

Add a comment

Comments

Plain text

  • Loading comments…