Emotional End-to-End Neural Speech Synthesizer
arXiv:1711.05447
Abstract
In this paper, we introduce an emotional speech synthesizer based on the recent end-to-end neural model, named Tacotron. Despite its benefits, we found that the original Tacotron suffers from the exposure bias problem and irregularity of the attention alignment. Later, we address the problem by utilization of context vector and residual connection at recurrent neural networks (RNNs). Our experiments showed that the model could successfully train and generate speech for given emotion labels.
5 pages, 3 figures
Cited by in corpus (5)
- Adversarial Training in Affective Computing and Sentiment Analysis: Recent Advances and Perspectives
- Voice Imitating Text-to-Speech Neural Networks
- MASS: Multi-task Anthropomorphic Speech Synthesis Framework
- ET-GAN: Cross-Language Emotion Transfer Based on Cycle-Consistent Generative Adversarial Networks
- A Fully Time-domain Neural Model for Subband-based Speech Synthesizer