12 papers
Let Me Look at You: Advanced Facial Expression Modeling for Conversational Speech Synthesis
Yifan Hu, Shuwei He, Rui Liu +1
Conversational Speech Synthesis is a fundamental component of human-computer interaction, aiming to generate contextually appropriate, expressive, and empathetic speech. However, f…
AuEmoChat: Authentic Emotion Understanding and Rendering for Conversational Speech Synthesis
Zhenqi Jia, Yuan Zhao, Aruukhan +2
Conversational Speech Synthesis (CSS) aims to synthesize speech with human-like emotional expression and contextual consistency in user-agent interactions. Existing CSS methods str…
AffectSpeech: A Large-Scale Emotional Speech Dataset with Fine-Grained Textual Descriptions for Speech Emotion Captioning and Synthesis
Tianhua Qi, Wenming Zheng, Björn W. Schuller +3
Emotion is essential in spoken communication, yet most existing frameworks in speech emotion modeling rely on predefined categories or low-dimensional continuous attributes, which…
Multimodal Fine-grained Context Interaction Graph Modeling for Conversational Speech Synthesis
Zhenqi Jia, Rui Liu, Berrak Sisman +1
Conversational Speech Synthesis (CSS) aims to generate speech with natural prosody by understanding the multimodal dialogue history (MDH). The latest work predicts the accurate pro…
NE-PADD: Leveraging Named Entity Knowledge for Robust Partial Audio Deepfake Detection via Attention Aggregation
Huhong Xian, Rui Liu, Berrak Sisman +1
Different from traditional sentence-level audio deepfake detection (ADD), partial audio deepfake detection (PADD) requires frame-level positioning of the location of fake speech. W…
UniTalker: Conversational Speech-Visual Synthesis
Yifan Hu, Rui Liu, Yi Ren +2
Conversational Speech Synthesis (CSS) is a key task in the user-agent interaction area, aiming to generate more expressive and empathetic speech for users. However, it is well-know…