Faces help you hear
Google's Looking to Listen separates one voice from a crowd by watching the speaker's face. Modalities work better when they explain each other.
Google Research published “Looking to Listen at the Cocktail Party” last week. You give it a video of several people talking over each other, pick a face, and it outputs just that person’s voice. The demos include two comedians shouting over each other and a crowded room, and the separation is startlingly clean.
The model takes two inputs. From the video, it runs a face model on each speaker’s face in every frame and extracts face embeddings. From the audio, it takes the spectrogram of the mixed soundtrack. The network combines them and predicts a mask for each speaker, which says for each frequency at each moment how much of that time-frequency cell belongs to that person. Apply the mask to the mixture and you get the isolated voice.
They trained it on a dataset they built from YouTube: about 100,000 clips of talks and lectures with one visible speaker and clean audio. They then mixed clips together synthetically, so the network always knew the right answer. That’s a nice trick. They didn’t need anyone to label who is speaking when, because they created the mixtures themselves.
Why does the face help so much? Because the audio alone is ambiguous. When two people with similar voices speak at once, a spectrogram doesn’t tell you which frequencies belong to whom. The face does: lips open and close in time with syllables, the jaw moves with loudness. The visual signal resolves exactly what the audio can’t. Humans do the same thing, and we’re much better at following one voice in a noisy bar when we can see the person’s mouth.
We’re building face technology at Amanda, so the application caught my attention. I’m especially interested in how the model combines sound and vision. Most multimodal systems I’ve seen treat modalities as parallel inputs: get features from the image, get features from the audio, concatenate, classify. This model uses one modality to disambiguate the other at a fine level of time and frequency. Vision tells audio where to look.
That’s how I’d expect intelligence in the physical world to work. Sound, sight, touch and motion aren’t separate streams that get summarized independently. They constantly explain each other. The sound of something breaking, the sight of it falling and the feel of it slipping out of your hand are all one event. A system that learns from all of them together should understand events much better than one trained on any single stream, and it can learn from the correlations between them without anyone labeling anything.