7 papers
UniDDT: Unifying Multimodal Understanding and Generation with Decoupled Diffusion Transformer
Shuai Wang, Liang Li, Yang Chen +3
Unified Multimodal Models (UMMs) have emerged as a critical direction for general-purpose multimodal intelligence, integrating understanding and generation into a single framework.…
Learning to Generate via Understanding: Understanding-Driven Intrinsic Rewarding for Unified Multimodal Models
Jiadong Pan, Liang Li, Yuxin Peng +6
Recently, unified multimodal models (UMMs) have made remarkable progress in integrating visual understanding and generation, demonstrating strong potential for complex text-to-imag…
InstructDubber: Instruction-based Alignment for Zero-shot Movie Dubbing
Zhedong Zhang, Liang Li, Gaoxiang Cong +5
Movie dubbing seeks to synthesize speech from a given script using a specific voice, while ensuring accurate lip synchronization and emotion-prosody alignment with the character's…
FlowDubber: Movie Dubbing with LLM-based Semantic-aware Learning and Flow Matching based Voice Enhancing
Gaoxiang Cong, Liang Li, Jiadong Pan +5
Movie Dubbing aims to convert scripts into speeches that align with the given movie clip in both temporal and emotional aspects while preserving the vocal timbre of a given brief r…
SafeCFG: Controlling Harmful Features with Dynamic Safe Guidance for Safe Generation
Jiadong Pan, Liang Li, Hongcheng Gao +3
Diffusion models (DMs) have demonstrated exceptional performance in text-to-image tasks, leading to their widespread use. With the introduction of classifier-free guidance (CFG), t…
EmoDubber: Towards High Quality and Emotion Controllable Movie Dubbing
Gaoxiang Cong, Jiadong Pan, Liang Li +5
Given a piece of text, a video clip, and a reference audio, the movie dubbing task aims to generate speech that aligns with the video while cloning the desired voice. The existing…