Emotion Intensity and its Control for Emotional Voice Conversion
arXiv:2201.03967 · doi:10.1109/TAFFC.2022.3175578
Abstract
Emotional voice conversion (EVC) seeks to convert the emotional state of an utterance while preserving the linguistic content and speaker identity. In EVC, emotions are usually treated as discrete categories overlooking the fact that speech also conveys emotions with various intensity levels that the listener can perceive. In this paper, we aim to explicitly characterize and control the intensity of emotion. We propose to disentangle the speaker style from linguistic content and encode the speaker style into a style embedding in a continuous space that forms the prototype of emotion embedding. We further learn the actual emotion encoder from an emotion-labelled database and study the use of relative attributes to represent fine-grained emotion intensity. To ensure emotional intelligibility, we incorporate emotion classification loss and emotion embedding similarity loss into the training of the EVC network. As desired, the proposed network controls the fine-grained emotion intensity in the output speech. Through both objective and subjective evaluations, we validate the effectiveness of the proposed network for emotional expressiveness and emotion intensity control.
Accepted by IEEE Transactions on Affective Computing
References in corpus (8)
- Parallel WaveNet: Fast High-Fidelity Speech Synthesis
- F0-consistent many-to-many non-parallel voice conversion via conditional autoencoder
- Any-to-Many Voice Conversion with Location-Relative Sequence-to-Sequence Modeling
- CHiVE: Varying Prosody in Speech Synthesis with a Linguistically Driven Dynamic Hierarchical Conditional Variational Network
- Fine-grained Style Modeling, Transfer and Prediction in Text-to-Speech Synthesis via Phone-Level Content-Style Disentanglement
- FastS2S-VC: Streaming Non-Autoregressive Sequence-to-Sequence Voice Conversion
- VAW-GAN for Disentanglement and Recomposition of Emotional Elements in Speech
- Recognition-Synthesis Based Non-Parallel Voice Conversion with Adversarial Learning
Cited by in corpus (6)
- An Overview of Affective Speech Synthesis and Conversion in the Deep Learning Era
- Probing Speech Emotion Recognition Transformers for Linguistic Knowledge
- Hierarchical Emotion Prediction and Control in Text-to-Speech Synthesis
- DurFlex-EVC: Duration-Flexible Emotional Voice Conversion Leveraging Discrete Representations without Text Alignment
- EmoSphere++: Emotion-Controllable Zero-Shot Text-to-Speech via Emotion-Adaptive Spherical Vector
- Attention-based Interactive Disentangling Network for Instance-level Emotional Voice Conversion