WaveNet: A Generative Model for Raw Audio
arXiv:1609.03499
Abstract
This paper introduces WaveNet, a deep neural network for generating raw audio waveforms. The model is fully probabilistic and autoregressive, with the predictive distribution for each audio sample conditioned on all previous ones; nonetheless we show that it can be efficiently trained on data with tens of thousands of samples per second of audio. When applied to text-to-speech, it yields state-of-the-art performance, with human listeners rating it as significantly more natural sounding than the best parametric and concatenative systems for both English and Mandarin. A single WaveNet can capture the characteristics of many different speakers with equal fidelity, and can switch between them by conditioning on the speaker identity. When trained to model music, we find that it generates novel and often highly realistic musical fragments. We also show that it can be employed as a discriminative model, returning promising results for phoneme recognition.
Cited by in corpus (36)
- Deep Learning for Audio Signal Processing
- Generating Long Sequences with Sparse Transformers
- The State of Sparsity in Deep Neural Networks
- AUTOVC: Zero-Shot Voice Style Transfer with Only Autoencoder Loss
- R-Transformer: Recurrent Neural Network Enhanced Transformer
- Waveform Modeling and Generation Using Hierarchical Recurrent Neural Networks for Speech Bandwidth Extension
- Generative adversarial network-based glottal waveform model for statistical parametric speech synthesis
- Assume, Augment and Learn: Unsupervised Few-Shot Meta-Learning via Random Labels and Data Augmentation
- Toward Interpretable Music Tagging with Self-Attention
- Adaptive Music Composition for Games
- The Indirect Convolution Algorithm
- Singing voice synthesis based on convolutional neural networks
- Learning to Speak Fluently in a Foreign Language: Multilingual Speech Synthesis and Cross-Language Voice Cloning
- GRASS: Generative Recursive Autoencoders for Shape Structures
- Rethinking Full Connectivity in Recurrent Neural Networks
- Non-Parallel Voice Conversion with Cyclic Variational Autoencoder
- Global Guarantees for Blind Demodulation with Generative Priors
- Recurrent Neural Networks with Stochastic Layers for Acoustic Novelty Detection
- Improving Image Classification Robustness through Selective CNN-Filters Fine-Tuning
- CNN Is All You Need
- Effective parameter estimation methods for an ExcitNet model in generative text-to-speech systems
- Accelerating Minibatch Stochastic Gradient Descent using Typicality Sampling
- A New GAN-based End-to-End TTS Training Algorithm
- Wasserstein-Wasserstein Auto-Encoders
- Hierarchical Sequence to Sequence Voice Conversion with Limited Data
- Equilibrated Recurrent Neural Network: Neuronal Time-Delayed Self-Feedback Improves Accuracy and Stability
- ASAC: Active Sensing using Actor-Critic models
- Speech denoising by parametric resynthesis
- Low-Latency Speaker-Independent Continuous Speech Separation
- Skeleton-Based Online Action Prediction Using Scale Selection Network
- Exploiting Syntactic Features in a Parsed Tree to Improve End-to-End TTS
- Voice command generation using Progressive Wavegans
- A Unified Neural Architecture for Instrumental Audio Tasks
- Scaling up deep neural networks: a capacity allocation perspective
- Semi-supervised and Population Based Training for Voice Commands Recognition
- Analyzing the benefits of communication channels between deep learning models