Pictures and words in the same space
OpenAI's CLIP learns from 400 million image and caption pairs to put images and text in one embedding space. Zero-shot classification is the demo. Shared embeddings are the real story.
OpenAI released two things last week. DALL·E generates images from text, and the samples (an armchair shaped like an avocado, a radish in a tutu walking a dog) got most of the attention. The other one, CLIP, is less fun to look at and I think more important.
CLIP is trained on 400 million pairs of images and text collected from the internet: a photo and its caption, alt text, or surrounding words. It has two encoders, one for images (a ResNet or a Vision Transformer) and one for text (a Transformer). Each maps its input into a vector in the same space. Training is contrastive: in each batch of, say, 32,000 pairs, the model is pushed to make each image’s vector close to its own caption’s vector and far from the other 31,999 captions. That’s the whole objective. It never predicts a label from a fixed list and it never generates anything.
What falls out is a shared space where images and text can be compared directly. To classify an image with no training on the task, you write the class names as sentences, “a photo of a dog,” “a photo of a cat,” encode them, and pick the one closest to the image. That zero-shot approach matches the accuracy of a ResNet-50 trained on ImageNet’s 1.28 million labeled examples, without seeing any of them. And it holds up much better than ImageNet models on sketches, cartoons and other shifted versions of the same objects.
Zero-shot ImageNet is the headline, but I think the shared space is the thing that will last. Once images and text live in one space, a lot of problems become nearest-neighbor search. Searching photos by any description. Finding the images that match a mood. Scoring whether a generated image matches its prompt, which is in fact how OpenAI reranks DALL·E’s samples. Matching content to what a user asks for in words.
I’ve been thinking about this from the recommendation side lately. Most recommenders know items by their IDs and learn what they are from who interacted with them. A shared embedding space where an item’s content, a user’s description of what they want and their past behavior are all comparable would change what a recommender can do, especially for new items with no history. You could describe what you’re in the mood for and get matched against the content itself.
The limits are the usual ones. The model learned from the internet, so it learned the internet’s biases, and OpenAI’s paper has a long section on that. It’s also bad at counting, spatial relations and anything that needs careful reasoning about an image. But as a general bridge between what things look like and how people describe them, it’s the best I’ve seen.