activity
20242026
collaborators

6 papers

cs.SD2026

CoSyncDiT: Cognitive Synchronous Diffusion Transformer for Movie Dubbing

Gaoxiang Cong, Liang Li, Jiaxin Ye +4

Movie dubbing aims to synthesize speech that preserves the vocal identity of a reference audio while synchronizing with the lip movements in a target video. Existing methods fail t…

cs.SD2025

InstructDubber: Instruction-based Alignment for Zero-shot Movie Dubbing

Zhedong Zhang, Liang Li, Gaoxiang Cong +5

Movie dubbing seeks to synthesize speech from a given script using a specific voice, while ensuring accurate lip synchronization and emotion-prosody alignment with the character's…

cs.MM2025

FlowDubber: Movie Dubbing with LLM-based Semantic-aware Learning and Flow Matching based Voice Enhancing

Gaoxiang Cong, Liang Li, Jiadong Pan +5

Movie Dubbing aims to convert scripts into speeches that align with the given movie clip in both temporal and emotional aspects while preserving the vocal timbre of a given brief r…

cs.SD2025

Prosody-Enhanced Acoustic Pre-training and Acoustic-Disentangled Prosody Adapting for Movie Dubbing

Zhedong Zhang, Liang Li, Chenggang Yan +3

Movie dubbing describes the process of transforming a script into speech that aligns temporally and emotionally with a given movie clip while exemplifying the speaker's voice demon…

cs.SD2024

Generating High-quality Symbolic Music Using Fine-grained Discriminators

Zhedong Zhang, Liang Li, Jiehua Zhang +5

Existing symbolic music generation methods usually utilize discriminator to improve the quality of generated music via global perception of music. However, considering the complexi…

cs.CL2024

StyleDubber: Towards Multi-Scale Style Learning for Movie Dubbing

Gaoxiang Cong, Yuankai Qi, Liang Li +6

Given a script, the challenge in Movie Dubbing (Visual Voice Cloning, V2C) is to generate speech that aligns well with the video in both time and emotion, based on the tone of a re…