A computer captioned a photo. It didn't see the photo
Google and Stanford both have networks that write sentences about images. How they work, and what they're actually learning.
Google published “Show and Tell” this week, and Karpathy and Fei-Fei Li at Stanford put out a similar paper around the same time. You give the model a photo and it writes a sentence, something like “A group of young people playing a game of frisbee.” Some of the examples in Google’s blog post are very good.
The architecture is simple, which is part of why I find it interesting. A convolutional network trained on ImageNet turns the image into a vector. That vector goes in as the first input to an LSTM, the same kind of recurrent network that’s been doing well on machine translation this year, and the LSTM generates the caption one word at a time, each word conditioned on the image and the words so far. The Google paper frames it literally as translation, from an image to English. It’s trained end to end on datasets like MSCOCO, where each image has several captions written by people.
What I keep coming back to is what “sees” means here. The image vector comes from a network trained to sort pictures into 1,000 categories, so it knows a lot about textures, object parts and typical scenes. The LSTM learns which word sequences tend to go with which image vectors. When it gets the frisbee caption right, a big part of the reason is that the training set had many photos of frisbee games that looked like this one. The error examples in the papers make this pretty clear. The wrong captions are usually plausible descriptions of a similar-looking scene, not descriptions of what’s actually in the photo.
I think it’s worth separating two claims. One is that the model maps images to the kind of sentence people write about them, which it clearly does, and which is already useful. Most photo search today runs on file names, tags and nearby text, so a model that writes a decent sentence for any photo is a big upgrade. The other claim is that the model understands the scene: that the frisbee is in the air and will come down, that the guy reaching for it will probably catch it or miss. Nothing in the training setup asks for that, and I’d be surprised if it’s in there.
This is a familiar gap in my own work. We can fit a 3D face to a photo quite accurately and still have no idea whether the person is about to laugh. The shape is in the image. What happens next is not in any single image.
I’d like to see the same kind of model trained on video with narration, where the sentence describes something happening over time. My guess is it would still mostly learn what events look like, unless the objective forces it to predict what comes next. Predicting the next word of a caption is a very different job from predicting the next second of the world.