works on

From the 1 of 5 linked papers with an AI index.

activity
20242026
collaborators

5 papers

cs.SD2026

AuEmoChat: Authentic Emotion Understanding and Rendering for Conversational Speech Synthesis

Zhenqi Jia, Yuan Zhao, Aruukhan +2

Conversational Speech Synthesis (CSS) aims to synthesize speech with human-like emotional expression and contextual consistency in user-agent interactions. Existing CSS methods str…

cs.MM2026

MAR3: Multi-Agent Recognition, Reasoning, and Reflection for Reference Audio-Visual Segmentation

Yuan Zhao, Zhenqi Jia, Yongqiang Zhang

The paper introduces MAR3, a training‑free multi‑agent framework that uses large language model agents to recognize expression difficulty, determine dominant modality, reason about…

cs.CL2025

Towards Authentic Movie Dubbing with Retrieve-Augmented Director-Actor Interaction Learning

Rui Liu, Yuan Zhao, Zhenqi Jia

The automatic movie dubbing model generates vivid speech from given scripts, replicating a speaker's timbre from a brief timbre prompt while ensuring lip-sync with the silent video…

cs.MM2024

Towards Expressive Video Dubbing with Multiscale Multimodal Context Interaction

Yuan Zhao, Rui Liu, Gaoxiang Cong

Automatic Video Dubbing (AVD) generates speech aligned with lip motion and facial emotion from scripts. Recent research focuses on modeling multimodal context to enhance prosody ex…

cs.MM2024

MCDubber: Multimodal Context-Aware Expressive Video Dubbing

Yuan Zhao, Zhenqi Jia, Rui Liu +3

Automatic Video Dubbing (AVD) aims to take the given script and generate speech that aligns with lip motion and prosody expressiveness. Current AVD models mainly utilize visual inf…