activity
20242026
collaborators

11 papers

cs.CV2026

EmotionDialogCN: A Spontaneous Multimodal Dataset for Mandarin Emotional Dialogue

Yi Zheng, Yifan Xu, Yan Zhou +8

Face-to-face audiovisual interaction is central to human communication, conveying rich emotional and social cues. However, existing multimodal dialogue datasets remain limited by i…

cs.SD2026

Speech Meets ELF: Audio Conditional Continuous-Target Diffusion for Speech Recognition and Translation

Xuanchen Li, Tianrui Wang, Yuheng Lu +11

Speech-to-text (S2T) systems for recognition (ASR) and translation (S2TT) typically generate discrete text tokens. In contrast, continuous-target language modelling performs genera…

cs.SD2026

Expressive Prompting: Improving Emotion Intensity and Speaker Consistency in Zero-Shot TTS

Haoyu Wang, Chunyu Qiang, Tianrui Wang +6

Recent advancements in speech synthesis have enabled large language model (LLM)-based systems to perform zero-shot generation with controllable content, timbre, speaker identity, a…

eess.AS2026

Efficient Emotion and Speaker Adaptation in LLM-Based TTS via Characteristic-Specific Partial Fine-Tuning

Tianrui Wang, Meng Ge, Cheng Gong +11

While LLM-based TTS models exhibit zero-shot emotion and speaker cloning, their cloning fidelity and pronunciation clarity degrade on unseen domains. Fine-tuning is essential for a…

eess.AS2026

Word-Level Emotional Expression Control in Zero-Shot Text-to-Speech Synthesis

Tianrui Wang, Haoyu Wang, Meng Ge +10

While emotional text-to-speech (TTS) has made significant progress, most existing research remains limited to utterance-level emotional expression and fails to support word-level c…

eess.AS2025

InstructAudio: Unified speech and music generation with natural language instruction

Chunyu Qiang, Kang Yin, Xiaopeng Wang +8

Text-to-speech (TTS) and text-to-music (TTM) models face significant limitations in instruction-based control. TTS systems usually depend on reference audio for timbre, offer only…