3 papers
cs.CV2025
MultiMotion: Multi Subject Video Motion Transfer via Video Diffusion Transformer
Penghui Liu, Jiangshan Wang, Yutong Shen +3
Multi-object video motion transfer poses significant challenges for Diffusion Transformer (DiT) architectures due to inherent motion entanglement and lack of object-level control.…
cs.CV2025
Follow-Your-Instruction: A Comprehensive MLLM Agent for World Data Synthesis
Kunyu Feng, Yue Ma, Xinhua Zhang +9
With the growing demands of AI-generated content (AIGC), the need for high-quality, diverse, and scalable data has become increasingly crucial. However, collecting large-scale real…
cs.CV2025
EVCtrl: Efficient Control Adapter for Visual Generation
Zixiang Yang, Yue Ma, Yinhan Zhang +3
Visual generation includes both image and video generation, training probabilistic models to create coherent, diverse, and semantically faithful content from scratch. While early r…