16 papers
Towards Compact Unified Multimodal Tracking: Synergizing Knowledge Distillation with Structural Pruning
Yuqi Li, Yuedong Tan, Huiran Duan +7
Unified multimodal object tracking has achieved remarkable robustness by leveraging complementary sensor data (e.g., RGB, Thermal, Depth), yet the heavy computational burden of sta…
Rethinking Layer-Wise Information Allocation for Vision Foundation Model Adaptation
Yuqi Li, Xi Xiao, Yunbei Zhang +6
Vision foundation models are increasingly reused as frozen backbones for downstream visual recognition, making parameter-efficient adaptation a central problem. Prompt-based adapta…
WorldCache: Accelerating World Models for Free via Heterogeneous Token Caching
Weilun Feng, Guoxin Fan, Haotong Qin +10
Diffusion-based world models have shown strong potential for unified world simulation, but the iterative denoising remains too costly for interactive use and long-horizon rollouts.…
Fast-SAM3D: 3Dfy Anything in Images but Faster
Weilun Feng, Mingqiang Wu, Zhiliang Chen +10
SAM3D enables scalable, open-world 3D reconstruction from complex scenes, yet its deployment is hindered by prohibitive inference latency. In this work, we conduct the \textbf{firs…
Echo-Forcing: A Scene Memory Framework for Interactive Long Video Generation
Mingqiang Wu, Weilun Feng, Zhefeng Zhang +8
Autoregressive video diffusion models enable open-ended generation through local attention and KV caching. However, existing training-free long-video optimization methods mainly fo…
Federated Knowledge Distillation for Multi-Model Architectures Lithography Hotspot Detection
Yuqi Li, Xingyou Lin, Yanli Li +6
As a special type of multimedia data, Lithography Hotspot Detection (LHD) training often requires stronger privacy protection than conventional multimedia data, and federated learn…