Efficient Neural Audio Synthesis
arXiv:1802.08435
Abstract
Sequential models achieve state-of-the-art results in audio, visual and textual domains with respect to both estimating the data distribution and generating high-quality samples. Efficient sampling for this class of models has however remained an elusive problem. With a focus on text-to-speech synthesis, we describe a set of general techniques for reducing sampling time while maintaining high output quality. We first describe a single-layer recurrent neural network, the WaveRNN, with a dual softmax layer that matches the quality of the state-of-the-art WaveNet model. The compact form of the network makes it possible to generate 24kHz 16-bit audio 4x faster than real time on a GPU. Second, we apply a weight pruning technique to reduce the number of weights in the WaveRNN. We find that, for a constant number of parameters, large sparse networks perform better than small dense networks and this relationship holds for sparsity levels beyond 96%. The small number of weights in a Sparse WaveRNN makes it possible to sample high-fidelity audio on a mobile CPU in real time. Finally, we propose a new generation scheme based on subscaling that folds a long sequence into a batch of shorter sequences and allows one to generate multiple samples at once. The Subscale WaveRNN produces 16 samples per step without loss of quality and offers an orthogonal method for increasing sampling efficiency.
10 pages
References in corpus (8)
- Empirical Evaluation of Gated Recurrent Neural Networks on Sequence Modeling
- Google's Neural Machine Translation System: Bridging the Gap between Human and Machine Translation
- WaveNet: A Generative Model for Raw Audio
- Non-Autoregressive Neural Machine Translation
- Deep Voice: Real-time Neural Text-to-Speech
- SampleRNN: An Unconditional End-to-End Neural Audio Generation Model
- Neural Machine Translation in Linear Time
- Block-Sparse Recurrent Neural Networks
Cited by in corpus (50)
- Glow-TTS: A Generative Flow for Text-to-Speech via Monotonic Alignment Search
- SuperNNova: an open-source framework for Bayesian, Neural Network based supernova classification
- End-to-end ASR: from Supervised to Semi-Supervised Learning with Modern Architectures
- MelNet: A Generative Model for Audio in the Frequency Domain
- Singing voice synthesis based on convolutional neural networks
- An Overview of Voice Conversion and its Challenges: From Statistical Modeling to Deep Learning
- Universal MelGAN: A Robust Neural Vocoder for High-Fidelity Waveform Generation in Multiple Domains
- Towards Robust Neural Vocoding for Speech Generation: A Survey
- TFGAN: Time and Frequency Domain Based Generative Adversarial Network for High-fidelity Speech Synthesis
- SqueezeWave: Extremely Lightweight Vocoders for On-device Speech Synthesis
- ConvS2S-VC: Fully convolutional sequence-to-sequence voice conversion
- A Spectral Energy Distance for Parallel Speech Synthesis
- WaveCycleGAN2: Time-domain Neural Post-filter for Speech Waveform Generation
- TTS-Portuguese Corpus: a corpus for speech synthesis in Brazilian Portuguese
- VARA-TTS: Non-Autoregressive Text-to-Speech Synthesis based on Very Deep VAE with Residual Attention
- Semi-Supervised Generative Modeling for Controllable Speech Synthesis
- Controllable neural text-to-speech synthesis using intuitive prosodic features
- FeatherWave: An efficient high-fidelity neural vocoder with multi-band linear prediction
- VoiceGrad: Non-Parallel Any-to-Many Voice Conversion with Annealed Langevin Dynamics
- Many-to-Many Voice Transformer Network
- DeviceTTS: A Small-Footprint, Fast, Stable Network for On-Device Text-to-Speech
- Towards achieving robust universal neural vocoding
- Singing Voice Synthesis Using Deep Autoregressive Neural Networks for Acoustic Modeling
- Latent Space Explorations of Singing Voice Synthesis using DDSP
- Interactive Text-to-Speech System via Joint Style Analysis
- Attention Forcing for Sequence-to-sequence Model Training
- Speech-to-Singing Conversion in an Encoder-Decoder Framework
- Capacity allocation analysis of neural networks: A tool for principled architecture design
- Synthesising Expressiveness in Peking Opera via Duration Informed Attention Network
- MoBoAligner: a Neural Alignment Model for Non-autoregressive TTS with Monotonic Boundary Search
- Improved Prosody from Learned F0 Codebook Representations for VQ-VAE Speech Waveform Reconstruction
- Noise Tokens: Learning Neural Noise Templates for Environment-Aware Speech Enhancement
- On-device neural speech synthesis
- High-Fidelity and Low-Latency Universal Neural Vocoder based on Multiband WaveRNN with Data-Driven Linear Prediction for Discrete Waveform Modeling
- Towards transformation-resilient provenance detection of digital media
- Nonparallel Voice Conversion with Augmented Classifier Star Generative Adversarial Networks
- FBWave: Efficient and Scalable Neural Vocoders for Streaming Text-To-Speech on the Edge
- Scyclone: High-Quality and Parallel-Data-Free Voice Conversion Using Spectrogram and Cycle-Consistent Adversarial Networks
- A unified sequence-to-sequence front-end model for Mandarin text-to-speech synthesis
- Capacity allocation through neural network layers
- PeriodNet: A non-autoregressive waveform generation model with a structure separating periodic and aperiodic components
- Unified Signal Compression Using Generative Adversarial Networks
- Audiovisual Speech Synthesis using Tacotron2
- Whispered and Lombard Neural Speech Synthesis
- Pretraining Strategies, Waveform Model Choice, and Acoustic Configurations for Multi-Speaker End-to-End Speech Synthesis
- Empirical Evaluation of Deep Learning Model Compression Techniques on the WaveNet Vocoder
- Low-Latency Real-Time Non-Parallel Voice Conversion based on Cyclic Variational Autoencoder and Multiband WaveRNN with Data-Driven Linear Prediction
- Enhancing Low-Quality Voice Recordings Using Disentangled Channel Factor and Neural Waveform Model
- Improving Accent Conversion with Reference Encoder and End-To-End Text-To-Speech
- Efficient And Scalable Neural Residual Waveform Coding With Collaborative Quantization