Expressive Speech Synthesis via Modeling Expressions with Variational Autoencoder
arXiv:1804.02135
Abstract
Recent advances in neural autoregressive models have improve the performance of speech synthesis (SS). However, as they lack the ability to model global characteristics of speech (such as speaker individualities or speaking styles), particularly when these characteristics have not been labeled, making neural autoregressive SS systems more expressive is still an open issue. In this paper, we propose to combine VoiceLoop, an autoregressive SS model, with Variational Autoencoder (VAE). This approach, unlike traditional autoregressive SS systems, uses VAE to model the global characteristics explicitly, enabling the expressiveness of the synthesized speech to be controlled in an unsupervised manner. Experiments using the VCTK and Blizzard2012 datasets show the VAE helps VoiceLoop to generate higher quality speech and to control the expressions in its synthesized speech by incorporating global characteristics into the speech generating process.
Accepted by Interspeech 2018
References in corpus (4)
Cited by in corpus (8)
- MelNet: A Generative Model for Audio in the Frequency Domain
- Deep Encoder-Decoder Models for Unsupervised Learning of Controllable Speech Synthesis
- Adversarial Training in Affective Computing and Sentiment Analysis: Recent Advances and Perspectives
- Modeling Prosodic Phrasing with Multi-Task Learning in Tacotron-based TTS
- Semi-Supervised Generative Modeling for Controllable Speech Synthesis
- Predicting Expressive Speaking Style From Text In End-To-End Speech Synthesis
- The Multi-speaker Multi-style Voice Cloning Challenge 2021
- In Other News: A Bi-style Text-to-speech Model for Synthesizing Newscaster Voice with Limited Data