Non-Attentive Tacotron: Robust and Controllable Neural TTS Synthesis Including Unsupervised Duration Modeling
arXiv:2010.04301
Abstract
This paper presents Non-Attentive Tacotron based on the Tacotron 2 text-to-speech model, replacing the attention mechanism with an explicit duration predictor. This improves robustness significantly as measured by unaligned duration ratio and word deletion rate, two metrics introduced in this paper for large-scale robustness evaluation using a pre-trained speech recognition model. With the use of Gaussian upsampling, Non-Attentive Tacotron achieves a 5-scale mean opinion score for naturalness of 4.41, slightly outperforming Tacotron 2. The duration predictor enables both utterance-wide and per-phoneme control of duration at inference time. When accurate target durations are scarce or unavailable in the training data, we propose a method using a fine-grained variational auto-encoder to train the duration predictor in a semi-supervised or unsupervised manner, with results almost as good as supervised training.
References in corpus (7)
- Attention-Based Models for Speech Recognition
- MelGAN: Generative Adversarial Networks for Conditional Waveform Synthesis
- CHiVE: Varying Prosody in Speech Synthesis with a Linguistically Driven Dynamic Hierarchical Conditional Variational Network
- AdaDurIAN: Few-shot Adaptation for Neural Text-to-Speech with DurIAN
- Parallel Tacotron: Non-Autoregressive and Controllable TTS
- JDI-T: Jointly trained Duration Informed Transformer for Text-To-Speech without Explicit Alignment
- TalkNet: Fully-Convolutional Non-Autoregressive Speech Synthesis Model
Cited by in corpus (23)
- A Survey on Neural Speech Synthesis
- Grad-TTS: A Diffusion Probabilistic Model for Text-to-Speech
- StyleTTS 2: Towards Human-Level Text-to-Speech through Style Diffusion and Adversarial Training with Large Speech Language Models
- Parallel Tacotron: Non-Autoregressive and Controllable TTS
- Daft-Exprt: Cross-Speaker Prosody Transfer on Any Text for Expressive Speech Synthesis
- Cross-lingual Low Resource Speaker Adaptation Using Phonological Features
- Neural HMMs are all you need (for high-quality attention-free TTS)
- Parallel Tacotron 2: A Non-Autoregressive Neural TTS Model with Differentiable Duration Modeling
- Creating New Voices using Normalizing Flows
- Investigating Content-Aware Neural Text-To-Speech MOS Prediction Using Prosodic and Linguistic Features
- Speech Synthesis and Control Using Differentiable DSP
- On granularity of prosodic representations in expressive text-to-speech
- Controllable speech synthesis by learning discrete phoneme-level prosodic representations
- Word-Level Style Control for Expressive, Non-attentive Speech Synthesis
- PnG BERT: Augmented BERT on Phonemes and Graphemes for Neural TTS
- WaveGrad 2: Iterative Refinement for Text-to-Speech Synthesis
- TalkNet 2: Non-Autoregressive Depth-Wise Separable Convolutional Model for Speech Synthesis with Explicit Pitch and Duration Prediction
- Fine-grained Noise Control for Multispeaker Speech Synthesis
- The Multi-speaker Multi-style Voice Cloning Challenge 2021
- A Neural TTS System with Parallel Prosody Transfer from Unseen Speakers
- VRAIN-UPV MLLP's system for the Blizzard Challenge 2021
- Improving Prosody for Unseen Texts in Speech Synthesis by Utilizing Linguistic Information and Noisy Data
- Preliminary study on using vector quantization latent spaces for TTS/VC systems with consistent performance