Vinson·Li

Essay No. 50

Read everything first, specialize later

OpenAI trained a Transformer to predict the next word on thousands of books, then fine-tuned it on small tasks. Language is getting its ImageNet moment.


OpenAI posted “Improving Language Understanding by Generative Pre-Training” last week. The method is simple and the results are strong. It’s the third paper this year pointing in the same direction, after ELMo from the Allen Institute and ULMFiT from fast.ai, so I think it’s time to call it a trend.

Their model is a 12-layer Transformer decoder, the attention-only architecture from last year, trained on about 7,000 unpublished books to predict the next word. That’s the “generative pre-training.” No labels, no task, just reading and guessing what comes next. Then they take that pretrained network and fine-tune it, with a small added output layer, on specific tasks: textual entailment, question answering, sentence similarity, classification. They report state-of-the-art results on 9 of the 12 benchmarks they tried, often by a lot, using the same model with minimal task-specific changes.

Why does predicting the next word help with question answering? Because to predict the next word well over thousands of books, a model has to pick up a lot: grammar, which words relate to which, how arguments and stories are structured, some facts about the world, some sense of what follows from what. All of that is useful for the downstream tasks. The labeled datasets for those tasks are small, sometimes a few thousand examples. They’re enough to teach the model the format of the task but not enough to teach it language.

This is the same thing that happened in computer vision around 2014. People stopped training image models from scratch and started with a network pretrained on ImageNet, then fine-tuned it. That made good vision models available to anyone with a few hundred labeled images. The difference is that language pretraining doesn’t need labels at all. Text itself is the supervision, and there’s an enormous amount of text.

My prediction: within a year or two, nobody will train a language model from scratch for a specific task. You’ll start from a big pretrained model and adapt it. And the pretrained models will get much bigger, because there’s no obvious limit on the data. This paper’s model is small by any standard and it still won most of its benchmarks.

The Transformer choice also seems to matter. The paper argues the attention architecture handles longer-range structure better than LSTMs do in transfer. Last summer I guessed Transformers would spread beyond translation. It took a year.

At Amanda we mostly work with images, not text, but I’m watching this closely. Pretraining on a lot of unlabeled data and adapting cheaply is a pattern that will reach every modality eventually.

Fin.

Add a comment

Comments

Plain text

  • Loading comments…