Vinson·Li

Essay No. 08

Two networks arguing

The GAN paper from NIPS this year: a generator, a discriminator, and a learned idea of what counts as real.


The paper I’ve reread most this month is “Generative Adversarial Nets,” by Ian Goodfellow and others at Montreal. It was on arXiv in June and presented at NIPS last week, and I completely missed it the first time around.

There are two networks. A generator takes random noise and outputs a sample, say an image. A discriminator takes an image and outputs the probability that it came from the real training data and not from the generator. The discriminator is trained to tell real from fake, and the generator is trained to make the discriminator wrong. The paper shows that if both have enough capacity and you could train them to the optimum, the generator ends up producing exactly the data distribution and the discriminator can do no better than a coin flip.

Most generative models I’ve seen write down an explicit probability for the data and try to maximize it, which becomes intractable very quickly for something like images. Here nobody writes down a probability. The discriminator is effectively a learned loss function. It tells the generator “this looks fake” in whatever sense it has figured out, and it keeps changing that sense as the generator gets better. That’s the part I find clever. The judge learns alongside the thing being judged.

The results in the paper are modest. MNIST digits, faces from the Toronto Face Database, CIFAR-10. The faces are small and a bit smeared and the CIFAR samples are mostly colored blobs. The authors are honest that training is delicate: if the generator gets too far ahead of the discriminator it can collapse onto a few outputs that happen to fool it, which they call “the Helvetica scenario.” Nobody would ship this.

For my own work it’s still interesting. Our faces come from fitting one template to a photo, so the output is always a deformation of the same mesh, and the quality is capped by how well I designed that deformation. When something looks wrong I find out by looking at it and wincing. A GAN-style system learns what faces look like from examples and generates new ones, and the “does this look real” judgment that I make by eye becomes a network. If it scales to high resolution, which is a big if, you could imagine generating faces and expressions directly.

The discriminator is also the part that makes me think past faces. For images, “real” is well defined because you have a dataset of real ones. For “is this a good photo” or “is this a good edit” there’s no fixed right answer, but you can collect examples of what people liked and train a critic on them. A generator trying to satisfy a learned critic of taste is an appealing idea for creative tools. This paper is the first time I’ve seen the basic loop work at all, even on blurry 28 by 28 digits.

Fin.

Add a comment

Comments

Plain text

  • Loading comments…