None of these people exist
Nvidia's StyleGAN generates photographic faces with control over pose, identity and freckles. A face company's view of faces becoming free.
Nvidia posted “A Style-Based Generator Architecture for Generative Adversarial Networks” last week, and the sample images are the most convincing fake faces I’ve ever seen. High resolution, varied ages and ethnicities, glasses, hats, hair falling across foreheads. None of these people exist.
Four years ago I wrote about the original GAN paper, whose faces were small and smeared. This is where four years of work went.
What’s new in StyleGAN is how the generator is organized. A normal GAN generator takes a random vector and upsamples it into an image. StyleGAN first passes the random vector through a mapping network, eight fully connected layers, to get an intermediate “style” vector. Then the image is built by a synthesis network starting from a learned constant, and the style vector is injected at every resolution through adaptive instance normalization, which scales and shifts the feature maps. Separately, random noise is added at each layer.
This gives a clean division of labor, which is the part I find beautiful. Styles injected at the coarse resolutions (4 by 4 up to 8 by 8) control pose, face shape, whether there are glasses. Middle resolutions control facial features and hairstyle. Fine resolutions control color scheme and microstructure. The per-layer noise handles stochastic details, like exactly where individual hairs fall, freckles, pores, which don’t change who the person is. You can take coarse styles from one generated face and fine styles from another and get a face that has the first’s pose and shape with the second’s coloring.
That’s close to what I tried to do by hand in my thesis and in the CAD paper: separate a face’s structure at different scales so you can control each one without disturbing the others. We built the hierarchy by hand from anatomy. StyleGAN learns one from 70,000 Flickr photos, and it lines up surprisingly well with what I would have designed.
For a face recognition company, this is both a curiosity and a warning. The good side: you can generate unlimited synthetic faces for testing and maybe training, with no privacy concerns because no real person is involved. The bad side: photos are no longer evidence that a person exists. Fake profiles, fake reviewers, fake ID photos are about to get very cheap. Anything that uses a single photo to prove identity will need liveness checks, depth, motion, or something the generator can’t fake as easily.
Nvidia says it will release the code and dataset. Once that happens, I’d expect generated faces all over the internet within months.