activity
20242026
most citedTowards Expressive Video Dubbing with Multiscale Multimodal Context Interaction

1 citations · 1 across the 3 of their papers we have counts for

collaborators

5 papers

cs.AI2026

STaR: Sensitive Trajectory Regulation for Unlearning in Large Reasoning Models

Jingjing Zhou, Gaoxiang Cong, Li Su +1

Large Reasoning Models (LRMs) have advanced automated multi-step reasoning, but their ability to generate complex Chain-of-Thought (CoT) trajectories introduces severe privacy risk…

cs.SD2025

InstructDubber: Instruction-based Alignment for Zero-shot Movie Dubbing

Zhedong Zhang, Liang Li, Gaoxiang Cong +5

Movie dubbing seeks to synthesize speech from a given script using a specific voice, while ensuring accurate lip synchronization and emotion-prosody alignment with the character's…

cs.MM2025

FlowDubber: Movie Dubbing with LLM-based Semantic-aware Learning and Flow Matching based Voice Enhancing

Gaoxiang Cong, Liang Li, Jiadong Pan +5

Movie Dubbing aims to convert scripts into speeches that align with the given movie clip in both temporal and emotional aspects while preserving the vocal timbre of a given brief r…

cs.MM20241 cited

Towards Expressive Video Dubbing with Multiscale Multimodal Context Interaction

Yuan Zhao, Rui Liu, Gaoxiang Cong

Automatic Video Dubbing (AVD) generates speech aligned with lip motion and facial emotion from scripts. Recent research focuses on modeling multimodal context to enhance prosody ex…

cs.SD2024

EmoDubber: Towards High Quality and Emotion Controllable Movie Dubbing

Gaoxiang Cong, Jiadong Pan, Liang Li +5

Given a piece of text, a video clip, and a reference audio, the movie dubbing task aims to generate speech that aligns with the video while cloning the desired voice. The existing…