GPT-4 reads the picture
OpenAI's GPT-4 scores near the top of the bar exam and can explain a joke in a photo. Describing a scene is useful, but I'm still unsure how far that gets it toward understanding the physical world.
OpenAI released GPT-4 on Tuesday. On the numbers they published, it’s a big step over GPT-3.5, the model behind ChatGPT: around the 90th percentile on a simulated bar exam where GPT-3.5 was around the 10th, strong scores on AP exams and the SAT, and noticeably better at following complex instructions. The technical report says almost nothing about the model itself. No parameter count, no architecture details, no training data description, citing competition and safety. That’s a real change from how this field has worked.
The demo that got the most attention was vision. GPT-4 accepts images as input, although that isn’t available to most users yet. In the livestream, Greg Brockman photographed a hand-drawn sketch of a website and GPT-4 wrote the HTML and JavaScript for it. In the report, it explains why an image of a phone plugged into a VGA connector is funny, and reads a chart to answer questions about it.
A few thoughts after using the text version for a few days.
It’s a lot better at reasoning through multi-step problems than ChatGPT was. It still makes things up, but less often and less blatantly. It’s the first model I’d trust to draft something complex that I then edit, as opposed to something I’d have to rewrite.
The image input is exciting, but I’d be careful about what it implies. Being able to describe an image, read text in it and answer questions about it is a lot of what CLIP-style models and captioning have been improving at for years, now combined with a very strong language model. That’s useful. It’s still understanding an image through language: what people would say about it.
What it doesn’t have is anything like physical experience. Ask it what happens if you push a glass off a table and it will give you a perfect answer, because that’s been written about a million times. But its knowledge that glass breaks comes from sentences about glass breaking, not from anything like having broken one. That works surprisingly far. I suspect it stops working in situations that aren’t described in text, which includes almost everything about fine physical interaction: how much force to use, how a specific cloth will fold, how to catch something you’ve never seen thrown.
I’m impressed, and I use it every day already. I’m also more convinced than ever that the next hard problem is grounding: connecting these models to the physical world through vision, action, touch and consequence, not only through descriptions of those things.