A Survey on Neural Speech Synthesis
arXiv:2106.15561
Abstract
Text to speech (TTS), or speech synthesis, which aims to synthesize intelligible and natural speech given text, is a hot research topic in speech, language, and machine learning communities and has broad applications in the industry. As the development of deep learning and artificial intelligence, neural network-based TTS has significantly improved the quality of synthesized speech in recent years. In this paper, we conduct a comprehensive survey on neural TTS, aiming to provide a good understanding of current research and future trends. We focus on the key components in neural TTS, including text analysis, acoustic models and vocoders, and several advanced topics, including fast TTS, low-resource TTS, robust TTS, expressive TTS, and adaptive TTS, etc. We further summarize resources related to TTS (e.g., datasets, opensource implementations) and discuss future research directions. This survey can serve both academic researchers and industry practitioners working on TTS.
A comprehensive survey on TTS, 63 pages, 18 tables, 7 figures, 457 references
References in corpus (54)
- Distilling the Knowledge in a Neural Network
- Neural Architecture Search with Reinforcement Learning
- Attention-Based Models for Speech Recognition
- NICE: Non-linear Independent Components Estimation
- MelGAN: Generative Adversarial Networks for Conditional Waveform Synthesis
- SampleRNN: An Unconditional End-to-End Neural Audio Generation Model
- MLS: A Large-Scale Multilingual Dataset for Speech Research
- Professor Forcing: A New Algorithm for Training Recurrent Networks
- Multilingual Neural Machine Translation with Knowledge Distillation
- Conditional Variational Autoencoder with Adversarial Learning for End-to-End Text-to-Speech
- MelNet: A Generative Model for Audio in the Frequency Domain
- JSUT corpus: free large-scale Japanese speech corpus for end-to-end speech synthesis
- Flowtron: an Autoregressive Flow-based Generative Network for Text-to-Speech Synthesis
- AdaSpeech: Adaptive Text to Speech for Custom Voice
- DDSP: Differentiable Digital Signal Processing
- RNN Approaches to Text Normalization: A Challenge
- On Fast Sampling of Diffusion Probabilistic Models
- NAS-BERT: Task-Agnostic and Adaptive-Size BERT Compression with Neural Architecture Search
- Learning to Efficiently Sample from Diffusion Probabilistic Models
- HiFiSinger: Towards High-Fidelity Neural Singing Voice Synthesis
- MultiSpeech: Multi-Speaker Text to Speech with Transformer
- What comprises a good talking-head video generation?: A Survey and Benchmark
- Towards Transfer Learning for End-to-End Speech Synthesis from Deep Pre-Trained Language Models
- Universal MelGAN: A Robust Neural Vocoder for High-Fidelity Waveform Generation in Multiple Domains
- AdaDurIAN: Few-shot Adaptation for Neural Text-to-Speech with DurIAN
- SqueezeWave: Extremely Lightweight Vocoders for On-device Speech Synthesis
- Review of end-to-end speech synthesis technology based on deep learning
- Cross-lingual Multi-speaker Text-to-speech Synthesis for Voice Cloning without Using Parallel Corpus for Unseen Speakers
- Expressive Neural Voice Cloning
- Model architectures to extrapolate emotional expressions in DNN-based text-to-speech
- VARA-TTS: Non-Autoregressive Text-to-Speech Synthesis based on Very Deep VAE with Residual Attention
- Unified Mandarin TTS Front-end Based on Distilled BERT Model
- Semi-Supervised Generative Modeling for Controllable Speech Synthesis
- TalkNet: Fully-Convolutional Non-Autoregressive Speech Synthesis Model
- AdaSpeech 3: Adaptive Text to Speech for Spontaneous Style
- Diff-TTS: A Denoising Diffusion Model for Text-to-Speech
- A Study of Multilingual Neural Machine Translation
- EfficientTTS: An Efficient and High-Quality Text-to-Speech Architecture
- Fre-GAN: Adversarial Frequency-consistent Audio Synthesis
- TTS-by-TTS: TTS-driven Data Augmentation for Fast and High-Quality Speech Synthesis
- LVCNet: Efficient Condition-Dependent Modeling Network for Waveform Generation
- Towards Multi-Scale Style Control for Expressive Speech Synthesis
- Exploring Machine Speech Chain for Domain Adaptation and Few-Shot Speaker Adaptation
- The Multi-speaker Multi-style Voice Cloning Challenge 2021
- s-Transformer: Segment-Transformer for Robust Neural Speech Synthesis
- Glow-WaveGAN: Learning Speech Representations from GAN-based Variational Auto-Encoder For High Fidelity Flow-based Speech Synthesis
- Universal Neural Vocoding with Parallel WaveNet
- Fast DCTTS: Efficient Deep Convolutional Text-to-Speech
- The AS-NU System for the M2VoC Challenge
- FeatherTTS: Robust and Efficient attention based Neural TTS
- Improve GAN-based Neural Vocoder using Pointwise Relativistic LeastSquare GAN
- Triple M: A Practical Text-to-speech Synthesis System With Multi-guidance Attention And Multi-band Multi-time LPCNet
- Improving Performance of Seen and Unseen Speech Style Transfer in End-to-end Neural TTS
- Pretraining Strategies, Waveform Model Choice, and Acoustic Configurations for Multi-Speaker End-to-End Speech Synthesis
Cited by in corpus (6)
- Augmented Datasheets for Speech Datasets and Ethical Decision-Making
- AdaSpeech 3: Adaptive Text to Speech for Spontaneous Style
- FastLTS: Non-Autoregressive End-to-End Unconstrained Lip-to-Speech Synthesis
- A Survey on Audio Synthesis and Audio-Visual Multimodal Processing
- fairseq S^2: A Scalable and Integrable Speech Synthesis Toolkit
- Improving Mandarin Prosodic Structure Prediction with Multi-level Contextual Information