Vinson·Li

Essay No. 29

16,000 samples a second

DeepMind's WaveNet generates raw audio one sample at a time, and its speech sounds far more human than anything before. How, and why it's so slow.


DeepMind published WaveNet last week, and if you haven’t listened to the samples on their blog, go do that first. The text-to-speech clips are the first computer voices I’ve heard that don’t sound like a computer. You can hear breaths, the mouth sounds between words, the small changes in pitch that make a voice sound like a person thinking about what they’re saying.

The approach is almost stubbornly direct. Most speech synthesis today either stitches together small recorded pieces of a real voice (concatenative) or generates parameters that a vocoder turns into sound (parametric). WaveNet does neither. It generates the raw waveform, one sample at a time, at 16,000 samples per second. Each sample is predicted from all the samples before it.

To make that possible, the audio is quantized to 256 levels with a mu-law curve, so predicting the next sample becomes a classification problem over 256 classes. The network is a stack of causal convolutions, meaning each output only looks at the past, and they’re dilated: each layer skips over more and more of the input, 1, 2, 4, 8 and so on up to 512 samples, then the pattern repeats. That lets the network see a fairly long window of past audio without needing thousands of layers. Condition it on text features and it speaks. Condition it on a speaker identity and it speaks in that voice. Train it on piano recordings and it produces something that sounds like a pianist improvising with no idea where the piece is going.

The obvious drawback is speed. Generating one second of audio means running the network 16,000 times in sequence, and each step depends on the one before, so you can’t parallelize across time. From what people are saying, the current version takes minutes of computation per second of speech. That’s fine for a research result and useless for a phone assistant.

I’m not worried about the speed. Someone will find a way to train a fast parallel model to imitate the slow one, or build hardware for it, and I’d guess WaveNet-quality voices ship in real products within a year or two.

What I find more significant is the general point. The field has been moving away from hand-designed intermediate representations. First features in vision, then phrase tables in translation, and now the vocoder parameters in speech. Every time, modeling the raw signal directly with enough data and a clever architecture won. Audio was the one I thought would hold out longest, because the signal is so long and so detailed. It didn’t.

Images, text and now sound can all be generated from scratch. Video is the obvious next one, and much harder, since you’d need to be consistent across both space and time. I’d guess it’s the last of the big ones.

Fin.

Add a comment

Comments

Plain text

  • Loading comments…