Deep Voice 2: Multi-Speaker Neural Text-to-Speech
arXiv:1705.08947
Abstract
We introduce a technique for augmenting neural text-to-speech (TTS) with lowdimensional trainable speaker embeddings to generate different voices from a single model. As a starting point, we show improvements over the two state-ofthe-art approaches for single-speaker neural TTS: Deep Voice 1 and Tacotron. We introduce Deep Voice 2, which is based on a similar pipeline with Deep Voice 1, but constructed with higher performance building blocks and demonstrates a significant audio quality improvement over Deep Voice 1. We improve Tacotron by introducing a post-processing neural vocoder, and demonstrate a significant audio quality improvement. We then demonstrate our technique for multi-speaker speech synthesis for both Deep Voice 2 and Tacotron on two multi-speaker TTS datasets. We show that a single neural TTS system can learn hundreds of unique voices from less than half an hour of data per speaker, while achieving high audio quality synthesis and preserving the speaker identities almost perfectly.
Accepted in NIPS 2017
References in corpus (9)
- Adam: A Method for Stochastic Optimization
- WaveNet: A Generative Model for Raw Audio
- Learning Phrase Representations using RNN Encoder-Decoder for Statistical Machine Translation
- Improved Techniques for Training GANs
- Deep Speaker: an End-to-End Neural Speaker Embedding System
- Deep Voice: Real-time Neural Text-to-Speech
- SampleRNN: An Unconditional End-to-End Neural Audio Generation Model
- Quasi-Recurrent Neural Networks
- Voice Conversion from Unaligned Corpora using Variational Autoencoding Wasserstein Generative Adversarial Networks
Cited by in corpus (26)
- Efficiently Trainable Text-to-Speech System Based on Deep Convolutional Networks with Guided Attention
- Towards End-to-End Prosody Transfer for Expressive Speech Synthesis with Tacotron
- Fast Spectrogram Inversion using Multi-head Convolutional Neural Networks
- An Overview of Affective Speech Synthesis and Conversion in the Deep Learning Era
- Flowtron: an Autoregressive Flow-based Generative Network for Text-to-Speech Synthesis
- NSML: A Machine Learning Platform That Enables You to Focus on Your Models
- Uncovering Latent Style Factors for Expressive Speech Synthesis
- Direct Speech-to-image Translation
- FloWaveNet : A Generative Flow for Raw Audio
- Highrisk Prediction from Electronic Medical Records via Deep Attention Networks
- TGAVC: Improving Autoencoder Voice Conversion with Text-Guided and Adversarial Training
- Voice Conversion with Conditional SampleRNN
- TTS-Portuguese Corpus: a corpus for speech synthesis in Brazilian Portuguese
- VSMask: Defending Against Voice Synthesis Attack via Real-Time Predictive Perturbation
- Emotional Prosody Control for Speech Generation
- MuSE-SVS: Multi-Singer Emotional Singing Voice Synthesizer that Controls Emotional Intensity
- Statistical Parametric Speech Synthesis Using Generative Adversarial Networks Under A Multi-task Learning Framework
- How does a spontaneously speaking conversational agent affect user behavior?
- Hierarchical Sequence to Sequence Voice Conversion with Limited Data
- Multi-task WaveNet: A Multi-task Generative Model for Statistical Parametric Speech Synthesis without Fundamental Frequency Conditions
- AlignTTS: Efficient Feed-Forward Text-to-Speech System without Explicit Alignment
- fairseq S^2: A Scalable and Integrable Speech Synthesis Toolkit
- Forward-Backward Decoding for Regularizing End-to-End TTS
- Reducing over-smoothness in speech synthesis using Generative Adversarial Networks
- Speech Rhythm-Based Speaker Embeddings Extraction from Phonemes and Phoneme Duration for Multi-Speaker Speech Synthesis
- Adversarially Trained Autoencoders for Parallel-Data-Free Voice Conversion