7 papers
Let Me Look at You: Advanced Facial Expression Modeling for Conversational Speech Synthesis
Yifan Hu, Shuwei He, Rui Liu +1
Conversational Speech Synthesis is a fundamental component of human-computer interaction, aiming to generate contextually appropriate, expressive, and empathetic speech. However, f…
FineCombo-TTS: Collaborative and Precise Controllable Speech Synthesis Using Text Descriptions and Reference Speech
Shuoyi Zhou, Yixuan Zhou, Peiji Yang +4
Controllable text-to-speech (TTS) has become a key research focus. However, methods based on either reference speech or text descriptions lack flexibility and precise control, and…
EmoTransCap: Dataset and Pipeline for Emotion Transition-Aware Speech Captioning in Discourses
Shuhao Xu, Yifan Hu, Jingjing Wu +3
Emotion perception and adaptive expression are fundamental capabilities in human-agent interaction. While recent advances in speech emotion captioning (SEC) have improved fine-grai…
TellWhisper: Tell Whisper Who Speaks When
Yifan Hu, Peiji Yang, Zhisheng Wang +2
Multi-speaker automatic speech recognition (MASR) aims to predict ''who spoke when and what'' from multi-speaker speech, a key technology for multi-party dialogue understanding. Ho…
UniTalker: Conversational Speech-Visual Synthesis
Yifan Hu, Rui Liu, Yi Ren +2
Conversational Speech Synthesis (CSS) is a key task in the user-agent interaction area, aiming to generate more expressive and empathetic speech for users. However, it is well-know…
Chain-Talker: Chain Understanding and Rendering for Empathetic Conversational Speech Synthesis
Yifan Hu, Rui Liu, Yi Ren +2
Conversational Speech Synthesis (CSS) aims to align synthesized speech with the emotional and stylistic context of user-agent interactions to achieve empathy. Current generative CS…