6 papers
EdgeFM: Efficient Edge Inference for Vision-Language Models
Mengling Deng, Yuanpeng Chen, Sheng Yang +12
Vision-language models (VLMs) have demonstrated strong applicability in edge industrial applications, yet their deployment remains severely constrained by requirements for determin…
EasyVFX: Frequency-Driven Decoupling for Resource-Efficient VFX Generation
Yue Ma, Xu Ye, Qinghe Wang +9
Generating high-fidelity visual effects (VFX) typically demands massive datasets and prohibitive computational power due to the intricate coupling of spatial textures and temporal…
Fast-BEV++: Fast by Algorithm, Deployable by Design
Yuanpeng Chen, Hui Song, Sheng Yang +5
The advancement of vision-only Bird's-Eye-View (BEV) perception, a core paradigm for cost-effective autonomous driving, is hindered by the long-standing fundamental trade-off betwe…
MultiMotion: Multi Subject Video Motion Transfer via Video Diffusion Transformer
Penghui Liu, Jiangshan Wang, Yutong Shen +3
Multi-object video motion transfer poses significant challenges for Diffusion Transformer (DiT) architectures due to inherent motion entanglement and lack of object-level control.…
EVCtrl: Efficient Control Adapter for Visual Generation
Zixiang Yang, Yue Ma, Yinhan Zhang +3
Visual generation includes both image and video generation, training probabilistic models to create coherent, diverse, and semantically faithful content from scratch. While early r…
Follow-Your-Instruction: A Comprehensive MLLM Agent for World Data Synthesis
Kunyu Feng, Yue Ma, Xinhua Zhang +9
With the growing demands of AI-generated content (AIGC), the need for high-quality, diverse, and scalable data has become increasingly crucial. However, collecting large-scale real…