5 papers
Differentiable Efficient Operator Search
Xiaohuan Pei, Jiyuan Zhang, Yuanfan Guo +4
Efficient multimodal foundation models often rely on manually designed token-reduction operators, such as pruning, merging, pooling, and adaptive reweighting. Although these operat…
Many-for-Many: Unify the Training of Multiple Video and Image Generation and Manipulation Tasks
Ruibin Li, Tao Yang, Yangming Shi +4
Diffusion models have shown impressive performance in many visual generation and manipulation tasks. Many existing methods focus on training a model for a specific task, especially…
Discriminator-Free Direct Preference Optimization for Video Diffusion
Haoran Cheng, Qide Dong, Liang Peng +7
Direct Preference Optimization (DPO), which aligns models with human preferences through win/lose data pairs, has achieved remarkable success in language and image generation. Howe…
LatentSync: Taming Audio-Conditioned Latent Diffusion Models for Lip Sync with SyncNet Supervision
Chunyu Li, Chao Zhang, Weikai Xu +6
End-to-end audio-conditioned latent diffusion models (LDMs) have been widely adopted for audio-driven portrait animation, demonstrating their effectiveness in generating lifelike a…
Text-to-Edit: Controllable End-to-End Video Ad Creation via Multimodal LLMs
Dabing Cheng, Haosen Zhan, Xingchen Zhao +6
The exponential growth of short-video content has ignited a surge in the necessity for efficient, automated solutions to video editing, with challenges arising from the need to und…