GPT-3 wrote a React component from a sentence
The demos flooding Twitter this week come from the same next-word objective as GPT-2, scaled a hundred times. Few-shot learning appeared on its own. Coding changes first.
My Twitter feed this week is mostly GPT-3 demos from people who got into OpenAI’s API beta. The one that got me was Sharif Shameem’s: he types a plain English description, like “a button that looks like a watermelon,” and the model writes the JSX for a React component that renders it. Others have it writing SQL from questions, generating guitar tabs, drafting emails, answering medical questions badly and confidently, and writing plausible short essays in someone else’s style.
The paper came out at the end of May. GPT-3 is the same basic design as GPT-2 from last year, a Transformer trained to predict the next word, with 175 billion parameters instead of 1.5 billion, trained on a much larger mix of web text, books and Wikipedia. There’s no new objective and no new architecture. Most of the paper is about what changes with scale.
The main finding is what they call in-context learning. You don’t fine-tune GPT-3 for a task. You write a few examples of the task into the prompt, like three English sentences with their French translations and then a fourth English sentence, and the model continues the pattern. That works across a surprising range of tasks, and it gets better with the number of examples and much better with model size. The smaller models barely do it at all. At 175 billion parameters it starts to rival fine-tuned models on some benchmarks.
When I wrote about GPT-2 last year, I said that if the scaling curve held for another factor of a hundred, the few-shot abilities wouldn’t stay a curiosity. It held.
I’m most interested in code, because it’s the application where I can judge quality myself and where I think the effect will show up first. Code is text with a very strong structure, the web is full of it, and a lot of everyday programming is exactly the pattern-continuation GPT-3 is good at: take this description and write the boilerplate that usually goes with it. The model doesn’t need to understand the whole system to save a developer twenty minutes on a form component.
It also fails in ways that should worry anyone using it. It’s fluent when it’s wrong. It invents functions that don’t exist. It can’t check its own work. For now I’d treat it like a very fast junior engineer who has read everything and verified nothing.
I wrote here in 2016 that conversational interfaces were at least ten years away. I’m now fairly sure I’ll be wrong about that, and early. Whatever makes these models useful for conversation won’t be a dialogue system someone designs. It’ll be a big model like this, trained more carefully, with some way to steer what it says.