Parallel WaveNet: Fast High-Fidelity Speech Synthesis
arXiv:1711.10433
Abstract
The recently-developed WaveNet architecture is the current state of the art in realistic speech synthesis, consistently rated as more natural sounding for many different languages than any previous system. However, because WaveNet relies on sequential generation of one audio sample at a time, it is poorly suited to today's massively parallel computers, and therefore hard to deploy in a real-time production setting. This paper introduces Probability Density Distillation, a new method for training a parallel feed-forward network from a trained WaveNet with no significant difference in quality. The resulting system is capable of generating high-fidelity speech samples at more than 20 times faster than real-time, and is deployed online by Google Assistant, including serving multiple English and Japanese voices.
Cited by in corpus (11)
- Physics-Constrained Deep Learning for High-dimensional Surrogate Modeling and Uncertainty Quantification without Labeled Data
- Deep Learning for Audio Signal Processing
- Computer-assisted Pronunciation Training -- Speech synthesis is almost all you need
- An Initial Investigation for Detecting Vocoder Fingerprints of Fake Audio
- Emotional Prosody Control for Speech Generation
- Non-parallel Voice Conversion System with WaveNet Vocoder and Collapsed Speech Suppression
- Spectrogram Inpainting for Interactive Generation of Instrument Sounds
- LVCNet: Efficient Condition-Dependent Modeling Network for Waveform Generation
- MelGlow: Efficient Waveform Generative Network Based on Location-Variable Convolution
- NeuralDPS: Neural Deterministic Plus Stochastic Model with Multiband Excitation for Noise-Controllable Waveform Generation
- Semi-supervised learning of glottal pulse positions in a neural analysis-synthesis framework