collaborators

5 papers

cs.LG2026

Differentiable Efficient Operator Search

Xiaohuan Pei, Jiyuan Zhang, Yuanfan Guo +4

Efficient multimodal foundation models often rely on manually designed token-reduction operators, such as pruning, merging, pooling, and adaptive reweighting. Although these operat…

cs.CV2026

Many-for-Many: Unify the Training of Multiple Video and Image Generation and Manipulation Tasks

Ruibin Li, Tao Yang, Yangming Shi +4

Diffusion models have shown impressive performance in many visual generation and manipulation tasks. Many existing methods focus on training a model for a specific task, especially…

cs.CV2025

Discriminator-Free Direct Preference Optimization for Video Diffusion

Haoran Cheng, Qide Dong, Liang Peng +7

Direct Preference Optimization (DPO), which aligns models with human preferences through win/lose data pairs, has achieved remarkable success in language and image generation. Howe…

cs.CV2025

LatentSync: Taming Audio-Conditioned Latent Diffusion Models for Lip Sync with SyncNet Supervision

Chunyu Li, Chao Zhang, Weikai Xu +6

End-to-end audio-conditioned latent diffusion models (LDMs) have been widely adopted for audio-driven portrait animation, demonstrating their effectiveness in generating lifelike a…

cs.CV2025

Text-to-Edit: Controllable End-to-End Video Ad Creation via Multimodal LLMs

Dabing Cheng, Haosen Zhan, Xingchen Zhao +6

The exponential growth of short-video content has ignited a surge in the necessity for efficient, automated solutions to video editing, with challenges arising from the need to und…