collaborators

7 papers

cs.SD2025

InstructDubber: Instruction-based Alignment for Zero-shot Movie Dubbing

Zhedong Zhang, Liang Li, Gaoxiang Cong +5

Movie dubbing seeks to synthesize speech from a given script using a specific voice, while ensuring accurate lip synchronization and emotion-prosody alignment with the character's…

cs.MM2025

FlowDubber: Movie Dubbing with LLM-based Semantic-aware Learning and Flow Matching based Voice Enhancing

Gaoxiang Cong, Liang Li, Jiadong Pan +5

Movie Dubbing aims to convert scripts into speeches that align with the given movie clip in both temporal and emotional aspects while preserving the vocal timbre of a given brief r…

cs.CV2025

ProgRoCC: A Progressive Approach to Rough Crowd Counting

Shengqin Jiang, Linfei Li, Haokui Zhang +6

As the number of individuals in a crowd grows, enumeration-based techniques become increasingly infeasible and their estimates increasingly unreliable. We propose instead an estima…

cs.CV2025

Visual and Semantic Prompt Collaboration for Generalized Zero-Shot Learning

Huajie Jiang, Zhengxian Li, Xiaohan Yu +4

Generalized zero-shot learning aims to recognize both seen and unseen classes with the help of semantic information that is shared among different classes. It inevitably requires c…

cs.CV2025

Collaborative Temporal Consistency Learning for Point-supervised Natural Language Video Localization

Zhuo Tao, Liang Li, Qi Chen +5

Natural language video localization (NLVL) is a crucial task in video understanding that aims to localize the target moment in videos specified by a given language description. Rec…

cs.SD2025

Prosody-Enhanced Acoustic Pre-training and Acoustic-Disentangled Prosody Adapting for Movie Dubbing

Zhedong Zhang, Liang Li, Chenggang Yan +3

Movie dubbing describes the process of transforming a script into speech that aligns temporally and emotionally with a given movie clip while exemplifying the speaker's voice demon…