Neural Codec Language Models are Zero-Shot Text to Speech Synthesizers
arXiv:2301.02111
Abstract
We introduce a language modeling approach for text to speech synthesis (TTS). Specifically, we train a neural codec language model (called Vall-E) using discrete codes derived from an off-the-shelf neural audio codec model, and regard TTS as a conditional language modeling task rather than continuous signal regression as in previous work. During the pre-training stage, we scale up the TTS training data to 60K hours of English speech which is hundreds of times larger than existing systems. Vall-E emerges in-context learning capabilities and can be used to synthesize high-quality personalized speech with only a 3-second enrolled recording of an unseen speaker as an acoustic prompt. Experiment results show that Vall-E significantly outperforms the state-of-the-art zero-shot TTS system in terms of speech naturalness and speaker similarity. In addition, we find Vall-E could preserve the speaker's emotion and acoustic environment of the acoustic prompt in synthesis. See https://aka.ms/valle for demos of our work.
Working in progress
Cited by in corpus (9)
- StyleTTS 2: Towards Human-Level Text-to-Speech through Style Diffusion and Adversarial Training with Large Speech Language Models
- UniCATS: A Unified Context-Aware Text-to-Speech Framework with Contextual VQ-Diffusion and Vocoding
- TextrolSpeech: A Text Style Control Speech Corpus With Codec Language Text-to-Speech Models
- USAT: A Universal Speaker-Adaptive Text-to-Speech Approach
- Speech Enhancement Using Continuous Embeddings of Neural Audio Codec
- Code-Mixed Text to Speech Synthesis under Low-Resource Constraints
- Zero-Shot vs. Few-Shot Multi-Speaker TTS Using Pre-trained Czech SpeechT5 Model
- Phase perturbation improves channel robustness for speech spoofing countermeasures
- An investigation of phrase break prediction in an End-to-End TTS system