Natively multimodal
Google announced Gemini, trained from the start on text, images, audio and video together. Why training on mixed modalities from day one matters more than bolting vision onto a language model.
Google announced Gemini last Wednesday, in three sizes: Ultra, Pro and Nano. Pro is now in Bard, Nano runs on the Pixel 8 Pro, and Ultra is coming early next year. I work at Google, but not on Gemini. I’m using the public blog posts and technical report here.
The benchmark numbers got the headlines, especially Ultra being the first model to score above human expert level on MMLU. The part I think matters more is in the first paragraph of the report: Gemini is “natively multimodal.” It was trained from the beginning on interleaved text, images, audio and video, rather than being a language model that learned to see later.
Most multimodal models so far have been assembled. You take a strong language model, take a separate vision encoder, often something like CLIP, and train a small adapter that converts image features into something the language model can read. It works well for describing images and answering questions about them. But the language model’s knowledge was built entirely from text, and vision is a guest in it. The model can only reason about images in the terms that the adapter translates them into.
A natively multimodal model learns all modalities in the same network at the same time. Images, audio and video frames go in as tokens in the same sequence as text, and the model learns relationships between them directly during pretraining. In principle it can learn things that are hard to express in text: the relationship between a sound and the motion that caused it, what changes between two video frames, how speech intonation changes meaning.
There’s been some criticism of one of the launch videos, which was edited and used still frames and text prompts rather than the real-time voice interaction it implied. Fair criticism. The technical report is more careful, and I’d judge the model by that and by using it.
Why I care: for years I’ve argued on this blog that understanding the physical world needs more than text. Video, sound, motion and eventually touch and action. Bolting vision onto a text model gets you a system that describes the world in words. Training on video and audio from the start is a step toward a system that learns from the world directly, where text is one modality among several rather than the foundation everything else is translated into.
It’s still a long way from a world model in the sense I care about. Gemini watches the world. It doesn’t act in it or feel the consequences. But the architecture, one model, many modalities, one sequence, is the one that can grow into that, and I’m glad the field is going this way.