Vinson·Li

Essay No. 120

GPT-4o hears you laugh

OpenAI and Google both showed real-time multimodal assistants this week. Speech as a native modality changes what a conversation with a model is, and latency turns out to be the feature.


Two keynotes a day apart. On Monday OpenAI announced GPT-4o, “o” for omni, a single model that takes text, audio and images as input and produces text, audio and images as output. On Tuesday at I/O, Google showed Project Astra, a prototype assistant that looks through your phone’s camera and talks with you about what it sees in real time, along with Veo, a new video generation model. I work at Google, not on either of these, and I’m going by the public demos.

GPT-4o’s voice demo is what everyone is talking about, and for good reason. Voice mode in ChatGPT until now was a pipeline: one model transcribes your speech into text, GPT-4 answers in text, and another model reads it aloud. OpenAI says that took several seconds, and everything that isn’t words got thrown away at the transcription step: tone, laughter, hesitation, who’s speaking, background sounds. GPT-4o is one network trained end to end across modalities, so the audio goes in as audio. OpenAI quotes an average response time around 320 milliseconds, which is in the range of human conversational turn-taking.

In the demos it laughs, changes its voice when asked to be more dramatic, sings, notices when someone is breathing too fast, and can be interrupted mid-sentence. Some of that is showmanship. But the latency and the ability to be interrupted are what make it feel like a conversation instead of a walkie-talkie.

It’s the same argument I made about Gemini being natively multimodal in December, applied to speech: bolting modalities together through text loses exactly the information that makes each modality valuable. A transcript of a sarcastic comment isn’t sarcastic. A model that hears the tone can respond to it.

Astra’s demo showed the other half. The assistant sees continuously through the camera, remembers what it saw a minute ago, and answers questions about the physical space: where did I leave my glasses, what does this part of the code on the screen do, what’s this neighborhood. It’s a step from a chatbot you consult toward an assistant that shares your context.

Put together, I think the week says that the interface to these models is going to be continuous, spoken, visual and real time, much more than text boxes. That changes products a lot. Latency becomes a first-class design constraint, the way it has always been for games. Memory of what happened in the conversation, and in the room, becomes part of the model’s job.

It’s also the closest thing yet to an AI that perceives the world the way we do, through sight and sound together, as it happens. It still doesn’t act in it or have a body, and I keep coming back to that. But the perception side of that picture is filling in faster than I expected.

Fin.

Add a comment

Comments

Plain text

  • Loading comments…