Deep Voice 3: Scaling Text-to-Speech with Convolutional Sequence Learning
arXiv:1710.07654
Abstract
We present Deep Voice 3, a fully-convolutional attention-based neural text-to-speech (TTS) system. Deep Voice 3 matches state-of-the-art neural speech synthesis systems in naturalness while training ten times faster. We scale Deep Voice 3 to data set sizes unprecedented for TTS, training on more than eight hundred hours of audio from over two thousand speakers. In addition, we identify common error modes of attention-based speech synthesis networks, demonstrate how to mitigate them, and compare several different waveform synthesis methods. We also describe how to scale inference to ten million queries per day on one single-GPU server.
Published as a conference paper at ICLR 2018. (v3 changed paper title)
References in corpus (7)
- Sequence to Sequence Learning with Neural Networks
- WaveNet: A Generative Model for Raw Audio
- Convolutional Sequence to Sequence Learning
- Attention-Based Models for Speech Recognition
- Deep Voice: Real-time Neural Text-to-Speech
- SampleRNN: An Unconditional End-to-End Neural Audio Generation Model
- Online and Linear-Time Attention by Enforcing Monotonic Alignments
Cited by in corpus (40)
- HiFi-GAN: Generative Adversarial Networks for Efficient and High Fidelity Speech Synthesis
- Style Tokens: Unsupervised Style Modeling, Control and Transfer in End-to-End Speech Synthesis
- Flowtron: an Autoregressive Flow-based Generative Network for Text-to-Speech Synthesis
- Semi-Supervised Neural Architecture Search
- Unsupervised End-to-End Learning of Discrete Linguistic Units for Voice Conversion
- MultiSpeech: Multi-Speaker Text to Speech with Transformer
- FloWaveNet : A Generative Flow for Raw Audio
- An Overview of Voice Conversion and its Challenges: From Statistical Modeling to Deep Learning
- Towards Robust Neural Vocoding for Speech Generation: A Survey
- Cross-lingual Multi-speaker Text-to-speech Synthesis for Voice Cloning without Using Parallel Corpus for Unseen Speakers
- Recent Advances and Trends in Multimodal Deep Learning: A Review
- ByteSing: A Chinese Singing Voice Synthesis System Using Duration Allocated Encoder-Decoder Acoustic Models and WaveRNN Vocoders
- TTS-Portuguese Corpus: a corpus for speech synthesis in Brazilian Portuguese
- Semi-Supervised Generative Modeling for Controllable Speech Synthesis
- VSMask: Defending Against Voice Synthesis Attack via Real-Time Predictive Perturbation
- Feature reinforcement with word embedding and parsing information in neural TTS
- Learning Individual Speaking Styles for Accurate Lip to Speech Synthesis
- Deep Text-to-Speech System with Seq2Seq Model
- Diff-TTS: A Denoising Diffusion Model for Text-to-Speech
- Noise Robust TTS for Low Resource Speakers using Pre-trained Model and Speech Enhancement
- Adversarially Trained Multi-Singer Sequence-To-Sequence Singing Synthesizer
- Hi-Fi Multi-Speaker English TTS Dataset
- WaveNODE: A Continuous Normalizing Flow for Speech Synthesis
- Controllable speech synthesis by learning discrete phoneme-level prosodic representations
- Hierarchical Sequence to Sequence Voice Conversion with Limited Data
- fairseq S^2: A Scalable and Integrable Speech Synthesis Toolkit
- TanhSoft -- a family of activation functions combining Tanh and Softplus
- Multi-task WaveNet: A Multi-task Generative Model for Statistical Parametric Speech Synthesis without Fundamental Frequency Conditions
- Building Multi lingual TTS using Cross Lingual Voice Conversion
- Code-Mixed Text to Speech Synthesis under Low-Resource Constraints
- Improved Prosodic Clustering for Multispeaker and Speaker-independent Phoneme-level Prosody Control
- Hierarchical Multi-Grained Generative Model for Expressive Speech Synthesis
- Multi-rate attention architecture for fast streamable Text-to-speech spectrum modeling
- Semi-supervised Learning for Multi-speaker Text-to-speech Synthesis Using Discrete Speech Representation
- Unsupervised Acoustic Unit Representation Learning for Voice Conversion using WaveNet Auto-encoders
- In Other News: A Bi-style Text-to-speech Model for Synthesizing Newscaster Voice with Limited Data
- Syntactic representation learning for neural network based TTS with syntactic parse tree traversal
- PeriodNet: A non-autoregressive waveform generation model with a structure separating periodic and aperiodic components
- Investigating context features hidden in End-to-End TTS
- Mathematical Vocoder Algorithm : Modified Spectral Inversion for Efficient Neural Speech Synthesis