WaveNet: A Generative Model for Raw Audio
arXiv:1609.03499
Abstract
This paper introduces WaveNet, a deep neural network for generating raw audio waveforms. The model is fully probabilistic and autoregressive, with the predictive distribution for each audio sample conditioned on all previous ones; nonetheless we show that it can be efficiently trained on data with tens of thousands of samples per second of audio. When applied to text-to-speech, it yields state-of-the-art performance, with human listeners rating it as significantly more natural sounding than the best parametric and concatenative systems for both English and Mandarin. A single WaveNet can capture the characteristics of many different speakers with equal fidelity, and can switch between them by conditioning on the speaker identity. When trained to model music, we find that it generates novel and often highly realistic musical fragments. We also show that it can be employed as a discriminative model, returning promising results for phoneme recognition.
Cited by in corpus (167)
- Deep Learning for Audio Signal Processing
- Towards Explainable Artificial Intelligence
- Generating Long Sequences with Sparse Transformers
- The State of Sparsity in Deep Neural Networks
- GANSynth: Adversarial Neural Audio Synthesis
- AUTOVC: Zero-Shot Voice Style Transfer with Only Autoencoder Loss
- Low Bit-Rate Speech Coding with VQ-VAE and a WaveNet Decoder
- F0-consistent many-to-many non-parallel voice conversion via conditional autoencoder
- R-Transformer: Recurrent Neural Network Enhanced Transformer
- Learning to dance: A graph convolutional adversarial network to generate realistic dance motions from audio
- Mining for Dark Matter Substructure: Inferring subhalo population properties from strong lenses with machine learning
- Detection of gravitational-wave signals from binary neutron star mergers using machine learning
- PolyGen: An Autoregressive Generative Model of 3D Meshes
- Waveform Modeling and Generation Using Hierarchical Recurrent Neural Networks for Speech Bandwidth Extension
- Expression Analysis Based on Face Regions in Read-world Conditions
- GeneraLight: Improving Environment Generalization of Traffic Signal Control via Meta Reinforcement Learning
- Deep Learning Ensemble for Real-time Gravitational Wave Detection of Spinning Binary Black Hole Mergers
- Generative adversarial network-based glottal waveform model for statistical parametric speech synthesis
- ATISS: Autoregressive Transformers for Indoor Scene Synthesis
- Toward Interpretable Music Tagging with Self-Attention
- Assume, Augment and Learn: Unsupervised Few-Shot Meta-Learning via Random Labels and Data Augmentation
- Adaptive Music Composition for Games
- Psychoacoustic Calibration of Loss Functions for Efficient End-to-End Neural Audio Coding
- Short-term daily precipitation forecasting with seasonally-integrated autoencoder
- The Indirect Convolution Algorithm
- Singing voice synthesis based on convolutional neural networks
- Learning to Speak Fluently in a Foreign Language: Multilingual Speech Synthesis and Cross-Language Voice Cloning
- GDP: Generalized Device Placement for Dataflow Graphs
- LAVARNET: Neural Network Modeling of Causal Variable Relationships for Multivariate Time Series Forecasting
- FastWave: Accelerating Autoregressive Convolutional Neural Networks on FPGA
- TFGAN: Time and Frequency Domain Based Generative Adversarial Network for High-fidelity Speech Synthesis
- GRASS: Generative Recursive Autoencoders for Shape Structures
- Cross-lingual Multi-speaker Text-to-speech Synthesis for Voice Cloning without Using Parallel Corpus for Unseen Speakers
- HiFi-GAN: High-Fidelity Denoising and Dereverberation Based on Speech Deep Features in Adversarial Networks
- Taming Visually Guided Sound Generation
- GAN-Aimbots: Using Machine Learning for Cheating in First Person Shooters
- Multi-Channel Auto-Calibration for the Atmospheric Imaging Assembly using Machine Learning
- Learning to Focus: Cascaded Feature Matching Network for Few-shot Image Recognition
- Unacceptable, where is my privacy? Exploring Accidental Triggers of Smart Speakers
- A Generative Model of Galactic Dust Emission Using Variational Inference
- Continuous Melody Generation via Disentangled Short-Term Representations and Structural Conditions
- Progressive Tandem Learning for Pattern Recognition with Deep Spiking Neural Networks
- Learning to Remember More with Less Memorization
- Rethinking Full Connectivity in Recurrent Neural Networks
- Feature reinforcement with word embedding and parsing information in neural TTS
- Non-Parallel Voice Conversion with Cyclic Variational Autoencoder
- Adversarial Generation of Time-Frequency Features with application in audio synthesis
- Online Training of Spiking Recurrent Neural Networks with Phase-Change Memory Synapses
- Adversarial Learning of Deepfakes in Accounting
- Hard-Coded Gaussian Attention for Neural Machine Translation
- Emotional Prosody Control for Speech Generation
- Cross-speaker Emotion Transfer Based on Speaker Condition Layer Normalization and Semi-Supervised Training in Text-To-Speech
- Multi-Target Emotional Voice Conversion With Neural Vocoders
- A non-causal FFTNet architecture for speech enhancement
- Phonetic Posteriorgrams based Many-to-Many Singing Voice Conversion via Adversarial Training
- Foley Music: Learning to Generate Music from Videos
- Neural Forecasting of the Italian Sovereign Bond Market with Economic News
- Learning Individual Speaking Styles for Accurate Lip to Speech Synthesis
- Global Guarantees for Blind Demodulation with Generative Priors
- Zero-Shot Multi-Speaker Text-To-Speech with State-of-the-art Neural Speaker Embeddings
- Recurrent Neural Networks with Stochastic Layers for Acoustic Novelty Detection
- FeatherWave: An efficient high-fidelity neural vocoder with multi-band linear prediction
- Image Transformation can make Neural Networks more robust against Adversarial Examples
- Improving Image Classification Robustness through Selective CNN-Filters Fine-Tuning
- Emotional Video to Audio Transformation Using Deep Recurrent Neural Networks and a Neuro-Fuzzy System
- Effective parameter estimation methods for an ExcitNet model in generative text-to-speech systems
- CNN Is All You Need
- Accelerating Minibatch Stochastic Gradient Descent using Typicality Sampling
- Voice Cloning: a Multi-Speaker Text-to-Speech Synthesis Approach based on Transfer Learning
- StarGAN-ZSVC: Towards Zero-Shot Voice Conversion in Low-Resource Contexts
- ConCare: Personalized Clinical Feature Embedding via Capturing the Healthcare Context
- Speech Synthesis and Control Using Differentiable DSP
- Music Artist Classification with Convolutional Recurrent Neural Networks
- A New GAN-based End-to-End TTS Training Algorithm
- Hierarchically Regularized Deep Forecasting
- A Brief Overview of Unsupervised Neural Speech Representation Learning
- Hierarchical Sequence to Sequence Voice Conversion with Limited Data
- Wasserstein-Wasserstein Auto-Encoders
- SPIN: A High Speed, High Resolution Vision Dataset for Tracking and Action Recognition in Ping Pong
- Vector-Quantized Timbre Representation
- Non-parallel Voice Conversion System with WaveNet Vocoder and Collapsed Speech Suppression
- Spectrogram Inpainting for Interactive Generation of Instrument Sounds
- AlignTTS: Efficient Feed-Forward Text-to-Speech System without Explicit Alignment
- LVCNet: Efficient Condition-Dependent Modeling Network for Waveform Generation
- UnivNet: A Neural Vocoder with Multi-Resolution Spectrogram Discriminators for High-Fidelity Waveform Generation
- A Study of Continuous Vector Representationsfor Theorem Proving
- AdaCare: Explainable Clinical Health Status Representation Learning via Scale-Adaptive Feature Extraction and Recalibration
- Diet deep generative audio models with structured lottery
- Speech denoising by parametric resynthesis
- Equilibrated Recurrent Neural Network: Neuronal Time-Delayed Self-Feedback Improves Accuracy and Stability
- A Novel Speech-Driven Lip-Sync Model with CNN and LSTM
- ASAC: Active Sensing using Actor-Critic models
- Multi-Task Time Series Forecasting With Shared Attention
- Label-efficient audio classification through multitask learning and self-supervision
- Learning compact generalizable neural representations supporting perceptual grouping
- Can Speaker Augmentation Improve Multi-Speaker End-to-End TTS?
- Alleviating Over-segmentation Errors by Detecting Action Boundaries
- On the Discrepancy between Density Estimation and Sequence Generation
- GraphTTS: graph-to-sequence modelling in neural text-to-speech
- DiffSVC: A Diffusion Probabilistic Model for Singing Voice Conversion
- Automated Respiratory Event Detection Using Deep Neural Networks
- Glow-WaveGAN: Learning Speech Representations from GAN-based Variational Auto-Encoder For High Fidelity Flow-based Speech Synthesis
- Knowledge-and-Data-Driven Amplitude Spectrum Prediction for Hierarchical Neural Vocoders
- Vision-Infused Deep Audio Inpainting
- Continuous Graph Flow
- Problems using deep generative models for probabilistic audio source separation
- Exploiting Syntactic Features in a Parsed Tree to Improve End-to-End TTS
- Unsupervised Acoustic Unit Representation Learning for Voice Conversion using WaveNet Auto-encoders
- You May Not Need Order in Time Series Forecasting
- Quasi-Periodic Parallel WaveGAN Vocoder: A Non-autoregressive Pitch-dependent Dilated Convolution Model for Parametric Speech Generation
- NeuralDPS: Neural Deterministic Plus Stochastic Model with Multiband Excitation for Noise-Controllable Waveform Generation
- The AS-NU System for the M2VoC Challenge
- Multi-rate attention architecture for fast streamable Text-to-speech spectrum modeling
- Universal Neural Vocoding with Parallel WaveNet
- Adversarial Robustness of Flow-Based Generative Models
- Hierarchical Recurrent Neural Networks for Conditional Melody Generation with Long-term Structure
- Go From the General to the Particular: Multi-Domain Translation with Domain Transformation Networks
- I'm Sorry for Your Loss: Spectrally-Based Audio Distances Are Bad at Pitch
- Uncertainty Intervals for Graph-based Spatio-Temporal Traffic Prediction
- MelGlow: Efficient Waveform Generative Network Based on Location-Variable Convolution
- Skeleton-Based Online Action Prediction Using Scale Selection Network
- On-device neural speech synthesis
- DarkGAN: Exploiting Knowledge Distillation for Comprehensible Audio Synthesis with GANs
- Low-Latency Speaker-Independent Continuous Speech Separation
- Short-time deep-learning based source separation for speech enhancement in reverberant environments with beamforming
- Detection and Evaluation of human and machine generated speech in spoofing attacks on automatic speaker verification systems
- FeatherTTS: Robust and Efficient attention based Neural TTS
- Multi-Scale Temporal Convolution Network for Classroom Voice Detection
- Pretraining Strategies, Waveform Model Choice, and Acoustic Configurations for Multi-Speaker End-to-End Speech Synthesis
- Semi-supervised learning of glottal pulse positions in a neural analysis-synthesis framework
- Prosody Learning Mechanism for Speech Synthesis System Without Text Length Limit
- Triple M: A Practical Text-to-speech Synthesis System With Multi-guidance Attention And Multi-band Multi-time LPCNet
- Autoencoding Neural Networks as Musical Audio Synthesizers
- Memory and attention in deep learning
- Basis-MelGAN: Efficient Neural Vocoder Based on Audio Decomposition
- Exploration of End-to-end Synthesisers forZero Resource Speech Challenge 2020
- Voice command generation using Progressive Wavegans
- Neural text-to-speech with a modeling-by-generation excitation vocoder
- Speech Representations and Phoneme Classification for Preserving the Endangered Language of Ladin
- Enhancing audio quality for expressive Neural Text-to-Speech
- AnyoneNet: Synchronized Speech and Talking Head Generation for Arbitrary Person
- Knowledge Distillation from BERT Transformer to Speech Transformer for Intent Classification
- A Deep-Bayesian Framework for Adaptive Speech Duration Modification
- JSSS: free Japanese speech corpus for summarization and simplification
- Recognition and Synthesis of Object Transport Motion
- Relational Data Selection for Data Augmentation of Speaker-dependent Multi-band MelGAN Vocoder
- Online Speaker Adaptation for WaveNet-based Neural Vocoders
- A Methodology for Exploring Deep Convolutional Features in Relation to Hand-Crafted Features with an Application to Music Audio Modeling
- Semi-supervised and Population Based Training for Voice Commands Recognition
- Granular Motor State Monitoring of Free Living Parkinson's Disease Patients via Deep Learning
- Efficient And Scalable Neural Residual Waveform Coding With Collaborative Quantization
- Leveraging Acoustic and Linguistic Embeddings from Pretrained speech and language Models for Intent Classification
- Automatic Feature Extraction for Heartbeat Anomaly Detection
- Analyzing the benefits of communication channels between deep learning models
- Real-time People Tracking and Identification from Sparse mm-Wave Radar Point-clouds
- Learned complex masks for multi-instrument source separation
- Denoising-and-Dereverberation Hierarchical Neural Vocoder for Robust Waveform Generation
- Effect of choice of probability distribution, randomness, and search methods for alignment modeling in sequence-to-sequence text-to-speech synthesis using hard alignment
- Sinusoidal wave generating network based on adversarial learning and its application: synthesizing frog sounds for data augmentation
- A Unified Neural Architecture for Instrumental Audio Tasks
- DeepA: A Deep Neural Analyzer For Speech And Singing Vocoding
- Revisiting Speech Content Privacy
- Scaling up deep neural networks: a capacity allocation perspective
- A Cyclical Post-filtering Approach to Mismatch Refinement of Neural Vocoder for Text-to-speech Systems
- GANtron: Emotional Speech Synthesis with Generative Adversarial Networks
- Efficient Modelling Across Time of Human Actions and Interactions
- Articulatory-WaveNet: Autoregressive Model For Acoustic-to-Articulatory Inversion