Vinson·Li

Essay No. 79

An image is worth 16x16 words

A paper under review at ICLR cuts images into patches and feeds them to a plain Transformer. With enough data, it beats convolutional networks. One architecture for everything is getting closer.


A paper submitted anonymously to ICLR last month has a title I’ve been repeating to people: “An Image is Worth 16x16 Words: Transformers for Image Recognition at Scale.” The authors aren’t officially named yet, but the JFT-300M dataset in the experiments is Google’s.

The method is straightforward. Cut the image into 16 by 16 pixel patches. Flatten each patch and project it with a linear layer to a vector. Add a learned position embedding so the model knows where each patch came from. Put a special classification token at the front. Feed the whole sequence into a standard Transformer encoder, the same architecture BERT uses for words, with no convolutions anywhere. Read the class from the output at the classification token.

A 224 by 224 image becomes a sentence of 196 “words,” and the Transformer does what it does with any sentence: every patch attends to every other patch, and the representations get refined layer by layer.

The finding that matters is about data. On ImageNet alone, this Vision Transformer does worse than a good ResNet of similar size. Convolutional networks have strong built-in assumptions, like nearby pixels being related and the same pattern meaning the same thing anywhere in the image. Those assumptions help a lot when data is limited. The Transformer doesn’t have them and has to learn them. But pretrained on JFT-300M, a dataset of 300 million images, the Transformer matches or beats the best convolutional networks, and it costs less compute to train to that level.

This is the Bitter Lesson again, in a very clean form. The built-in structure of convolutions was right about images, and it helped for years. With enough data, a general architecture learns that structure, and apparently some other useful things too, and the built-in version becomes a limit.

What excites me more is the direction. Three years ago Transformers were a translation model. Then they took over language with GPT and BERT. Now images. The same architecture is showing up in speech, protein structure and recommendation. When one architecture works on text, images and audio, you can start mixing them: feed image patches and words into the same model, in the same sequence, and let attention relate them directly. I’d guess we’ll see large models trained on images and text together within a year, and that the boundary between “vision model” and “language model” will start to blur. Video would follow, as a sequence of patches across space and time, although the length of that sequence is a real problem at the moment.

Fin.

Add a comment

Comments

Plain text

  • Loading comments…