High Quality Streaming Speech Synthesis with Low, Sentence-Length-Independent Latency
arXiv:2111.09052 · doi:10.21437/Interspeech.2020-2464
Abstract
This paper presents an end-to-end text-to-speech system with low latency on a CPU, suitable for real-time applications. The system is composed of an autoregressive attention-based sequence-to-sequence acoustic model and the LPCNet vocoder for waveform generation. An acoustic model architecture that adopts modules from both the Tacotron 1 and 2 models is proposed, while stability is ensured by using a recently proposed purely location-based attention mechanism, suitable for arbitrary sentence length generation. During inference, the decoder is unrolled and acoustic feature generation is performed in a streaming manner, allowing for a nearly constant latency which is independent from the sentence length. Experimental results show that the acoustic model can produce feature sequences with minimal latency about 31 times faster than real-time on a computer CPU and 6.5 times on a mobile CPU, enabling it to meet the conditions required for real-time applications on both devices. The full end-to-end system can generate almost natural quality speech, which is verified by listening tests.
Proceedings of INTERSPEECH 2020
References in corpus (6)
- PyTorch: An Imperative Style, High-Performance Deep Learning Library
- WaveNet: A Generative Model for Raw Audio
- Attention-Based Models for Speech Recognition
- MelGAN: Generative Adversarial Networks for Conditional Waveform Synthesis
- Parallel WaveNet: Fast High-Fidelity Speech Synthesis
- MelNet: A Generative Model for Audio in the Frequency Domain
Cited by in corpus (13)
- A Survey on Neural Speech Synthesis
- Review of end-to-end speech synthesis technology based on deep learning
- Cross-lingual Low Resource Speaker Adaptation Using Phonological Features
- LightGrad: Lightweight Diffusion Probabilistic Model for Text-to-Speech
- Investigating Content-Aware Neural Text-To-Speech MOS Prediction Using Prosodic and Linguistic Features
- Controllable speech synthesis by learning discrete phoneme-level prosodic representations
- Prosodic Clustering for Phoneme-level Prosody Control in End-to-End Speech Synthesis
- Fine-grained Noise Control for Multispeaker Speech Synthesis
- Word-Level Style Control for Expressive, Non-attentive Speech Synthesis
- Karaoker: Alignment-free singing voice synthesis with speech training data
- Improved Prosodic Clustering for Multispeaker and Speaker-independent Phoneme-level Prosody Control
- Rapping-Singing Voice Synthesis based on Phoneme-level Prosody Control
- Low-Latency Incremental Text-to-Speech Synthesis with Distilled Context Prediction Network