10 papers
Goku: A Million-Scale Universal Dataset and Benchmark for Instruction-Based Video Editing
Sen Liang, Cong Wang, Zhentao Yu +8
Existing instruction-based video editing datasets commonly focus on single-task appearance editing, failing to meet the complex creative demands of real-world scenarios. To bridge…
CineDance: Towards Next-Generation Multi-Shot Long-Form Cinematic Audio-Video Generation
Yuheng Chen, Teng Hu, Yuji Wang +7
The fidelity and structural diversity of training datasets fundamentally determine the capabilities of video generation models. While commercial systems showremarkableabilitytogene…
Omni-Customizer: End-to-End MultiModal Customization for Joint Audio-Video Generation
Yuheng Chen, Qingdong He, Teng Hu +4
The landscape of joint audio and video generation has been fundamentally transformed by the advent of powerful foundation models. Despite these strides, achieving cohesive multimod…
UniAVGen: Unified Audio and Video Generation with Asymmetric Cross-Modal Interactions
Guozhen Zhang, Zixiang Zhou, Teng Hu +6
Due to the lack of effective cross-modal modeling, existing open-source audio-video generation methods often exhibit compromised lip synchronization and insufficient semantic consi…
Making Avatars Interact: Towards Text-Driven Human-Object Interaction for Controllable Talking Avatars
Youliang Zhang, Zhengguang Zhou, Zhentao Yu +11
Generating talking avatars is a fundamental task in video generation. Although existing methods can generate full-body talking avatars with simple human motion, extending this task…
InstanceV: Instance-Level Video Generation
Yuheng Chen, Teng Hu, Jiangning Zhang +3
Recent advances in text-to-video diffusion models have enabled the generation of high-quality videos conditioned on textual descriptions. However, most existing text-to-video model…