Van Gogh is a Gram matrix
A new paper separates the content of an image from its style using a network trained for classification. How it works, and what it suggests about taste.
Gatys, Ecker and Bethge from Tübingen posted “A Neural Algorithm of Artistic Style” last week. Give it a photo and a painting, and it produces the photo painted in the style of the painting. A picture of houses along a river in Tübingen, done as Starry Night, as Munch’s The Scream, as a Kandinsky. The results are much better than any style filter I’ve seen.
The method uses a VGG network trained on ImageNet classification, nothing trained for art at all. They define two things.
Content is the feature activations at some higher layer. If two images produce similar activations there, they contain similar things in similar places, roughly the same houses and the same river.
Style is where it gets interesting. At each layer they take the feature maps and compute the Gram matrix, which is the correlation between every pair of feature channels, summed over all positions in the image. That throws away where things are and keeps only which features tend to appear together. If “short curved stroke” and “saturated yellow” usually fire together in Starry Night, that co-occurrence ends up in the Gram matrix, regardless of whether it’s in the sky or the village.
Then they start from noise and optimize the pixels, the same trick as DeepDream, until the image’s content features match the photo and its Gram matrices at several layers match the painting. The result has the photo’s layout and the painting’s texture statistics.
I like how concrete this makes something I’d have called ineffable. Style, at least the part of it this captures, is which low-level visual features co-occur, independent of layout. It’s statistics about brushstrokes and colors. It doesn’t capture everything. Starry Night isn’t great only because of its brushstrokes, and the method has no idea why Van Gogh composed the scene the way he did. But the surface of a style turns out to be measurable with a network that was never shown a painting during training.
I keep wondering where this goes. If style can be written as statistics, you can measure distances between styles, search by style, and maybe learn which statistics people find beautiful. That’s a first step toward a machine that has some kind of taste, even if a shallow one. I’d love to see someone collect human ratings of images and train on them directly. The ratings would be noisy and people would disagree, but the disagreement is information too.