From the 1 of 5 linked papers with an AI index.
5 papers
AuEmoChat: Authentic Emotion Understanding and Rendering for Conversational Speech Synthesis
Zhenqi Jia, Yuan Zhao, Aruukhan +2
Conversational Speech Synthesis (CSS) aims to synthesize speech with human-like emotional expression and contextual consistency in user-agent interactions. Existing CSS methods str…
MAR3: Multi-Agent Recognition, Reasoning, and Reflection for Reference Audio-Visual Segmentation
Yuan Zhao, Zhenqi Jia, Yongqiang Zhang
The paper introduces MAR3, a training‑free multi‑agent framework that uses large language model agents to recognize expression difficulty, determine dominant modality, reason about…
Towards Authentic Movie Dubbing with Retrieve-Augmented Director-Actor Interaction Learning
Rui Liu, Yuan Zhao, Zhenqi Jia
The automatic movie dubbing model generates vivid speech from given scripts, replicating a speaker's timbre from a brief timbre prompt while ensuring lip-sync with the silent video…
Towards Expressive Video Dubbing with Multiscale Multimodal Context Interaction
Yuan Zhao, Rui Liu, Gaoxiang Cong
Automatic Video Dubbing (AVD) generates speech aligned with lip motion and facial emotion from scripts. Recent research focuses on modeling multimodal context to enhance prosody ex…
MCDubber: Multimodal Context-Aware Expressive Video Dubbing
Yuan Zhao, Zhenqi Jia, Rui Liu +3
Automatic Video Dubbing (AVD) aims to take the given script and generate speech that aligns with lip motion and prosody expressiveness. Current AVD models mainly utilize visual inf…