FastSpeech 2: Fast and High-Quality End-to-End Text to Speech
arXiv:2006.04558
Abstract
Non-autoregressive text to speech (TTS) models such as FastSpeech can synthesize speech significantly faster than previous autoregressive models with comparable quality. The training of FastSpeech model relies on an autoregressive teacher model for duration prediction (to provide more information as input) and knowledge distillation (to simplify the data distribution in output), which can ease the one-to-many mapping problem (i.e., multiple speech variations correspond to the same text) in TTS. However, FastSpeech has several disadvantages: 1) the teacher-student distillation pipeline is complicated and time-consuming, 2) the duration extracted from the teacher model is not accurate enough, and the target mel-spectrograms distilled from teacher model suffer from information loss due to data simplification, both of which limit the voice quality. In this paper, we propose FastSpeech 2, which addresses the issues in FastSpeech and better solves the one-to-many mapping problem in TTS by 1) directly training the model with ground-truth target instead of the simplified output from teacher, and 2) introducing more variation information of speech (e.g., pitch, energy and more accurate duration) as conditional inputs. Specifically, we extract duration, pitch and energy from speech waveform and directly take them as conditional inputs in training and use predicted values in inference. We further design FastSpeech 2s, which is the first attempt to directly generate speech waveform from text in parallel, enjoying the benefit of fully end-to-end inference. Experimental results show that 1) FastSpeech 2 achieves a 3x training speed-up over FastSpeech, and FastSpeech 2s enjoys even faster inference speed; 2) FastSpeech 2 and 2s outperform FastSpeech in voice quality, and FastSpeech 2 can even surpass autoregressive models. Audio samples are available at https://speechresearch.github.io/fastspeech2/.
Accepted by ICLR 2021
References in corpus (8)
- WaveNet: A Generative Model for Raw Audio
- MelGAN: Generative Adversarial Networks for Conditional Waveform Synthesis
- Deep Voice: Real-time Neural Text-to-Speech
- Glow-TTS: A Generative Flow for Text-to-Speech via Monotonic Alignment Search
- AdaSpeech: Adaptive Text to Speech for Custom Voice
- DDSP: Differentiable Digital Signal Processing
- End-to-End Adversarial Text-to-Speech
- JDI-T: Jointly trained Duration Informed Transformer for Text-To-Speech without Explicit Alignment
Cited by in corpus (63)
- An Overview of Affective Speech Synthesis and Conversion in the Deep Learning Era
- VQTTS: High-Fidelity Text-to-Speech Synthesis with Self-Supervised VQ Acoustic Feature
- HiFiSinger: Towards High-Fidelity Neural Singing Voice Synthesis
- Transfer Learning Framework for Low-Resource Text-to-Speech using a Large-Scale Unlabeled Speech Corpus
- Fine-grained Style Modeling, Transfer and Prediction in Text-to-Speech Synthesis via Phone-Level Content-Style Disentanglement
- UniCATS: A Unified Context-Aware Text-to-Speech Framework with Contextual VQ-Diffusion and Vocoding
- WavThruVec: Latent speech representation as intermediate features for neural speech synthesis
- Daft-Exprt: Cross-Speaker Prosody Transfer on Any Text for Expressive Speech Synthesis
- Review of end-to-end speech synthesis technology based on deep learning
- TGAVC: Improving Autoencoder Voice Conversion with Text-Guided and Adversarial Training
- TextrolSpeech: A Text Style Control Speech Corpus With Codec Language Text-to-Speech Models
- StyleSinger: Style Transfer for Out-of-Domain Singing Voice Synthesis
- SQuId: Measuring Speech Naturalness in Many Languages
- Heterogeneous Target Speech Separation
- SANE-TTS: Stable And Natural End-to-End Multilingual Text-to-Speech
- Hierarchical Emotion Prediction and Control in Text-to-Speech Synthesis
- Text-Free Prosody-Aware Generative Spoken Language Modeling
- Prosody-controllable spontaneous TTS with neural HMMs
- TTS-Guided Training for Accent Conversion Without Parallel Data
- Rich Prosody Diversity Modelling with Phone-level Mixture Density Network
- LightGrad: Lightweight Diffusion Probabilistic Model for Text-to-Speech
- Emotional Prosody Control for Speech Generation
- Forensic Analysis and Localization of Multiply Compressed MP3 Audio Using Transformers
- BiFSMN: Binary Neural Network for Keyword Spotting
- Creating New Voices using Normalizing Flows
- MuSE-SVS: Multi-Singer Emotional Singing Voice Synthesizer that Controls Emotional Intensity
- Zero-shot text-to-speech synthesis conditioned using self-supervised speech representation model
- FastLTS: Non-Autoregressive End-to-End Unconstrained Lip-to-Speech Synthesis
- SNAC: Speaker-normalized affine coupling layer in flow-based architecture for zero-shot multi-speaker text-to-speech
- Unsupervised TTS Acoustic Modeling for TTS with Conditional Disentangled Sequential VAE
- EfficientTTS: An Efficient and High-Quality Text-to-Speech Architecture
- ContextSpeech: Expressive and Efficient Text-to-Speech for Paragraph Reading
- Parallel Synthesis for Autoregressive Speech Generation
- On granularity of prosodic representations in expressive text-to-speech
- On the Impact of Voice Anonymization on Speech Diagnostic Applications: a Case Study on COVID-19 Detection
- Multi-stage Progressive Compression of Conformer Transducer for On-device Speech Recognition
- A Survey on Audio Synthesis and Audio-Visual Multimodal Processing
- FlexLip: A Controllable Text-to-Lip System
- How does a spontaneously speaking conversational agent affect user behavior?
- Fine-grained Noise Control for Multispeaker Speech Synthesis
- fairseq S^2: A Scalable and Integrable Speech Synthesis Toolkit
- DeepSinger: Singing Voice Synthesis with Data Mined From the Web
- LightSpeech: Lightweight and Fast Text to Speech with Neural Architecture Search
- Code-Mixed Text to Speech Synthesis under Low-Resource Constraints
- Style Equalization: Unsupervised Learning of Controllable Generative Sequence Models
- Energy-Based Models with Applications to Speech and Language Processing
- FastLR: Non-Autoregressive Lipreading Model with Integrate-and-Fire
- The Multi-speaker Multi-style Voice Cloning Challenge 2021
- Multi-Singer: Fast Multi-Singer Singing Voice Vocoder With A Large-Scale Corpus
- Speech Rhythm-Based Speaker Embeddings Extraction from Phonemes and Phoneme Duration for Multi-Speaker Speech Synthesis
- Using IPA-Based Tacotron for Data Efficient Cross-Lingual Speaker Adaptation and Pronunciation Enhancement
- Visual-Aware Text-to-Speech
- Adversarial Learning of Intermediate Acoustic Feature for End-to-End Lightweight Text-to-Speech
- Indonesian-English Code-Switching Speech Synthesizer Utilizing Multilingual STEN-TTS and Bert LID
- Controllable Generation of Artificial Speaker Embeddings through Discovery of Principal Directions
- Varianceflow: High-Quality and Controllable Text-to-Speech using Variance Information via Normalizing Flow
- Rhythm-controllable Attention with High Robustness for Long Sentence Speech Synthesis
- An Overview on Generative AI at Scale with Edge-Cloud Computing
- AdaVocoder: Adaptive Vocoder for Custom Voice
- An Empirical Study on End-to-End Singing Voice Synthesis with Encoder-Decoder Architectures
- Pruning Self-Attention for Zero-Shot Multi-Speaker Text-to-Speech
- FADEL: Uncertainty-aware Fake Audio Detection with Evidential Deep Learning
- Emotion-Coherent Speech Data Augmentation and Self-Supervised Contrastive Style Training for Enhancing Kids's Story Speech Synthesis