On audio
- Veo 3 has sound, and that changes the medium
Google's Veo 3 generates video with synchronized dialogue, sound effects and ambient audio. Sound turns clips into scenes. Also at I/O: a language model that writes by denoising.
2 min reads likes comments - Text to music is a representation problem
Google Research's MusicLM generates music from text descriptions. The interesting part is its stack of tokens: one for meaning, one for sound, one shared between music and words.
2 min reads likes comments - Faces help you hear
Google's Looking to Listen separates one voice from a crowd by watching the speaker's face. Modalities work better when they explain each other.
2 min reads likes comments - 16,000 samples a second
DeepMind's WaveNet generates raw audio one sample at a time, and its speech sounds far more human than anything before. How, and why it's so slow.
2 min reads likes comments