On multimodal
- GPT-4o hears you laugh
OpenAI and Google both showed real-time multimodal assistants this week. Speech as a native modality changes what a conversation with a model is, and latency turns out to be the feature.
2 min reads likes comments - Natively multimodal
Google announced Gemini, trained from the start on text, images, audio and video together. Why training on mixed modalities from day one matters more than bolting vision onto a language model.
2 min reads likes comments - GPT-4 reads the picture
OpenAI's GPT-4 scores near the top of the bar exam and can explain a joke in a photo. Describing a scene is useful, but I'm still unsure how far that gets it toward understanding the physical world.
2 min reads likes comments - One architecture, any input
DeepMind's Perceiver handles images, audio, video and point clouds with the same network, by cross-attending to a small latent array. Modalities stop needing their own models.
2 min reads likes comments - Pictures and words in the same space
OpenAI's CLIP learns from 400 million image and caption pairs to put images and text in one embedding space. Zero-shot classification is the demo. Shared embeddings are the real story.
2 min reads likes comments - Faces help you hear
Google's Looking to Listen separates one voice from a crowd by watching the speaker's face. Modalities work better when they explain each other.
2 min reads likes comments