Emotional Voice Conversion: Theory, Databases and ESD
arXiv:2105.14762
Abstract
In this paper, we first provide a review of the state-of-the-art emotional voice conversion research, and the existing emotional speech databases. We then motivate the development of a novel emotional speech database (ESD) that addresses the increasing research need. With this paper, the ESD database is now made available to the research community. The ESD database consists of 350 parallel utterances spoken by 10 native English and 10 native Chinese speakers and covers 5 emotion categories (neutral, happy, angry, sad and surprise). More than 29 hours of speech data were recorded in a controlled acoustic environment. The database is suitable for multi-speaker and cross-lingual emotional voice conversion studies. As case studies, we implement several state-of-the-art emotional voice conversion systems on the ESD database. This paper provides a reference study on ESD in conjunction with its release.
Speech Communication
References in corpus (12)
- Parallel-Data-Free Voice Conversion Using Cycle-Consistent Adversarial Networks
- Voice Transformer Network: Sequence-to-Sequence Voice Conversion Using Transformer with Text-to-Speech Pretraining
- An Overview of Voice Conversion and its Challenges: From Statistical Modeling to Deep Learning
- Expressive TTS Training with Frame and Style Reconstruction Loss
- Reinforcement Learning Based Emotional Editing Constraint Conversation Generation
- Converting Anyone's Emotion: Towards Speaker-Independent Emotional Voice Conversion
- Multi-Target Emotional Voice Conversion With Neural Vocoders
- VAW-GAN for Singing Voice Conversion with Non-parallel Training Data
- Joint training framework for text-to-speech and voice conversion using multi-source Tacotron and WaveNet
- Reinforcement Learning for Emotional Text-to-Speech Synthesis with Improved Emotion Discriminability
- Limited Data Emotional Voice Conversion Leveraging Text-to-Speech: Two-stage Sequence-to-Sequence Training
- ICE-Talk: an Interface for a Controllable Expressive Talking Machine