An Overview of Affective Speech Synthesis and Conversion in the Deep Learning Era
arXiv:2210.03538 · doi:10.1109/JPROC.2023.3250266
Abstract
Speech is the fundamental mode of human communication, and its synthesis has long been a core priority in human-computer interaction research. In recent years, machines have managed to master the art of generating speech that is understandable by humans. But the linguistic content of an utterance encompasses only a part of its meaning. Affect, or expressivity, has the capacity to turn speech into a medium capable of conveying intimate thoughts, feelings, and emotions -- aspects that are essential for engaging and naturalistic interpersonal communication. While the goal of imparting expressivity to synthesised utterances has so far remained elusive, following recent advances in text-to-speech synthesis, a paradigm shift is well under way in the fields of affective speech synthesis and conversion as well. Deep learning, as the technology which underlies most of the recent advances in artificial intelligence, is spearheading these efforts. In the present overview, we outline ongoing trends and summarise state-of-the-art approaches in an attempt to provide a comprehensive overview of this exciting field.
Submitted to the Proceedings of IEEE
References in corpus (13)
- WaveNet: A Generative Model for Raw Audio
- HiFi-GAN: Generative Adversarial Networks for Efficient and High Fidelity Speech Synthesis
- MelGAN: Generative Adversarial Networks for Conditional Waveform Synthesis
- Deep Voice: Real-time Neural Text-to-Speech
- Parallel WaveNet: Fast High-Fidelity Speech Synthesis
- A Survey on Neural Speech Synthesis
- Emotion Intensity and its Control for Emotional Voice Conversion
- Probing Speech Emotion Recognition Transformers for Linguistic Knowledge
- Emotional Voice Conversion With Cycle-consistent Adversarial Network
- The ICML 2022 Expressive Vocalizations Workshop and Competition: Recognizing, Generating, and Personalizing Vocal Bursts
- StarGAN-based Emotional Voice Conversion for Japanese Phrases
- Textless Speech Emotion Conversion using Discrete and Decomposed Representations
- Generating Diverse Vocal Bursts with StyleGAN2 and MEL-Spectrograms
Cited by in corpus (9)
- Not My Voice! A Taxonomy of Ethical and Safety Harms of Speech Generators
- Leveraging AI-Generated Emotional Self-Voice to Nudge People towards their Ideal Selves
- Hierarchical Emotion Prediction and Control in Text-to-Speech Synthesis
- DurFlex-EVC: Duration-Flexible Emotional Voice Conversion Leveraging Discrete Representations without Text Alignment
- Exploring the Potential of Metacognitive Support Agents for Human-AI Co-Creation
- SpeakEasy: Enhancing Text-to-Speech Interactions for Expressive Content Creation
- StyleSpeech: Parameter-efficient Fine Tuning for Pre-trained Controllable Text-to-Speech
- DRKF: Decoupled Representations with Knowledge Fusion for Multimodal Emotion Recognition
- Effect of Attention and Self-Supervised Speech Embeddings on Non-Semantic Speech Tasks