FastSpeech: Fast, Robust and Controllable Text to Speech
arXiv:1905.09263
Abstract
Neural network based end-to-end text to speech (TTS) has significantly improved the quality of synthesized speech. Prominent methods (e.g., Tacotron 2) usually first generate mel-spectrogram from text, and then synthesize speech from the mel-spectrogram using vocoder such as WaveNet. Compared with traditional concatenative and statistical parametric approaches, neural network based end-to-end models suffer from slow inference speed, and the synthesized speech is usually not robust (i.e., some words are skipped or repeated) and lack of controllability (voice speed or prosody control). In this work, we propose a novel feed-forward network based on Transformer to generate mel-spectrogram in parallel for TTS. Specifically, we extract attention alignments from an encoder-decoder based teacher model for phoneme duration prediction, which is used by a length regulator to expand the source phoneme sequence to match the length of the target mel-spectrogram sequence for parallel mel-spectrogram generation. Experiments on the LJSpeech dataset show that our parallel model matches autoregressive models in terms of speech quality, nearly eliminates the problem of word skipping and repeating in particularly hard cases, and can adjust voice speed smoothly. Most importantly, compared with autoregressive Transformer TTS, our model speeds up mel-spectrogram generation by 270x and the end-to-end speech synthesis by 38x. Therefore, we call our model FastSpeech.
Accepted by NeurIPS2019
References in corpus (3)
Cited by in corpus (71)
- A Comparative Study on Transformer vs RNN in Speech Applications
- Conditional Variational Autoencoder with Adversarial Learning for End-to-End Text-to-Speech
- DiffWave: A Versatile Diffusion Model for Audio Synthesis
- Any-to-Many Voice Conversion with Location-Relative Sequence-to-Sequence Modeling
- An Overview of Affective Speech Synthesis and Conversion in the Deep Learning Era
- WaveFlow: A Compact Flow-based Model for Raw Audio
- Text-Conditioned Transformer for Automatic Pronunciation Error Detection
- Data Augmentation for End-to-end Code-switching Speech Recognition
- KazakhTTS: An Open-Source Kazakh Text-to-Speech Synthesis Dataset
- High Quality Streaming Speech Synthesis with Low, Sentence-Length-Independent Latency
- Transfer Learning Framework for Low-Resource Text-to-Speech using a Large-Scale Unlabeled Speech Corpus
- VoiceFixer: Toward General Speech Restoration with Neural Vocoder
- Fine-grained Style Modeling, Transfer and Prediction in Text-to-Speech Synthesis via Phone-Level Content-Style Disentanglement
- Controllable Data Generation by Deep Learning: A Review
- TSNAT: Two-Step Non-Autoregressvie Transformer Models for Speech Recognition
- Towards Robust Neural Vocoding for Speech Generation: A Survey
- WavThruVec: Latent speech representation as intermediate features for neural speech synthesis
- Giving Commands to a Self-Driving Car: How to Deal with Uncertain Situations?
- Parallel Tacotron: Non-Autoregressive and Controllable TTS
- XiaoiceSing: A High-Quality and Integrated Singing Voice Synthesis System
- TGAVC: Improving Autoencoder Voice Conversion with Text-Guided and Adversarial Training
- Heterogeneous Target Speech Separation
- SANE-TTS: Stable And Natural End-to-End Multilingual Text-to-Speech
- Semi-Supervised Generative Modeling for Controllable Speech Synthesis
- Token-Level Ensemble Distillation for Grapheme-to-Phoneme Conversion
- Rich Prosody Diversity Modelling with Phone-level Mixture Density Network
- Learning to Maximize Speech Quality Directly Using MOS Prediction for Neural Text-to-Speech
- LightGrad: Lightweight Diffusion Probabilistic Model for Text-to-Speech
- Autovocoder: Fast Waveform Generation from a Learned Speech Representation using Differentiable Digital Signal Processing
- Parallel Tacotron 2: A Non-Autoregressive Neural TTS Model with Differentiable Duration Modeling
- Emotional Prosody Control for Speech Generation
- Creating New Voices using Normalizing Flows
- MuSE-SVS: Multi-Singer Emotional Singing Voice Synthesizer that Controls Emotional Intensity
- FastLTS: Non-Autoregressive End-to-End Unconstrained Lip-to-Speech Synthesis
- Diff-TTS: A Denoising Diffusion Model for Text-to-Speech
- SNAC: Speaker-normalized affine coupling layer in flow-based architecture for zero-shot multi-speaker text-to-speech
- Unsupervised TTS Acoustic Modeling for TTS with Conditional Disentangled Sequential VAE
- Fine-Tuning by Curriculum Learning for Non-Autoregressive Neural Machine Translation
- Attention based end to end Speech Recognition for Voice Search in Hindi and English
- WaveNODE: A Continuous Normalizing Flow for Speech Synthesis
- ContextSpeech: Expressive and Efficient Text-to-Speech for Paragraph Reading
- Controllable speech synthesis by learning discrete phoneme-level prosodic representations
- Incremental Text-to-Speech Synthesis with Prefix-to-Prefix Framework
- Probing the phonetic and phonological knowledge of tones in Mandarin TTS models
- A Survey on Audio Synthesis and Audio-Visual Multimodal Processing
- DeepSinger: Singing Voice Synthesis with Data Mined From the Web
- Non-autoregressive End-to-end Speech Translation with Parallel Autoregressive Rescoring
- Exploring Machine Speech Chain for Domain Adaptation and Few-Shot Speaker Adaptation
- AlignTTS: Efficient Feed-Forward Text-to-Speech System without Explicit Alignment
- Adapting an Unadaptable ASR System
- DiDiSpeech: A Large Scale Mandarin Speech Corpus
- Towards Natural Bilingual and Code-Switched Speech Synthesis Based on Mix of Monolingual Recordings and Cross-Lingual Voice Conversion
- The Multi-speaker Multi-style Voice Cloning Challenge 2021
- Task-Level Curriculum Learning for Non-Autoregressive Neural Machine Translation
- MelGlow: Efficient Waveform Generative Network Based on Location-Variable Convolution
- Transferring Source Style in Non-Parallel Voice Conversion
- Rapping-Singing Voice Synthesis based on Phoneme-level Prosody Control
- DelightfulTTS: The Microsoft Speech Synthesis System for Blizzard Challenge 2021
- FCEM: A Novel Fast Correlation Extract Model For Real Time Steganalysis of VoIP Stream via Multi-head Attention
- PTeacher: a Computer-Aided Personalized Pronunciation Training System with Exaggerated Audio-Visual Corrective Feedback
- Varianceflow: High-Quality and Controllable Text-to-Speech using Variance Information via Normalizing Flow
- Towards Natural and Controllable Cross-Lingual Voice Conversion Based on Neural TTS Model and Phonetic Posteriorgram
- An Overview on Generative AI at Scale with Edge-Cloud Computing
- Adversarial Learning of Intermediate Acoustic Feature for End-to-End Lightweight Text-to-Speech
- High-Quality Real Time Facial Capture Based on Single Camera
- Improving Prosody for Unseen Texts in Speech Synthesis by Utilizing Linguistic Information and Noisy Data
- AnyoneNet: Synchronized Speech and Talking Head Generation for Arbitrary Person
- Hierarchical Context-Aware Transformers for Non-Autoregressive Text to Speech
- Zero-Shot Text-to-Speech for Text-Based Insertion in Audio Narration
- RefineGAN: Universally Generating Waveform Better than Ground Truth with Highly Accurate Pitch and Intensity Responses
- Effect of choice of probability distribution, randomness, and search methods for alignment modeling in sequence-to-sequence text-to-speech synthesis using hard alignment