12 papers
Towards Compact Unified Multimodal Tracking: Synergizing Knowledge Distillation with Structural Pruning
Yuqi Li, Yuedong Tan, Huiran Duan +7
Unified multimodal object tracking has achieved remarkable robustness by leveraging complementary sensor data (e.g., RGB, Thermal, Depth), yet the heavy computational burden of sta…
RSICCLLM: A Multimodal Large Language Model for Remote Sensing Image Change Captioning
Yelin Wang, Zijia Song, Shuo Ye +6
Remote Sensing Image Change Captioning (RSICC) aims to describe changes between bi-temporal remote sensing images and holds significant research and application value. However, mos…
DFM: Difference Feature Modeling with Text-Guided Gated Contrastive Loss for Remote Sensing Image Change Captioning
Yelin Wang, Zijia Song, Chuanguang Yang +4
The primary goal of Remote Sensing Image Change Captioning (RSICC) is to automatically generate descriptions of changes between remote sensing images captured at different time poi…
Fast-SAM3D: 3Dfy Anything in Images but Faster
Weilun Feng, Mingqiang Wu, Zhiliang Chen +10
SAM3D enables scalable, open-world 3D reconstruction from complex scenes, yet its deployment is hindered by prohibitive inference latency. In this work, we conduct the \textbf{firs…
Not All Tasks Quantize Equally: Fisher-Guided Quantization for Visual Geometry Transformer
Yipu Zhang, Jintao Cheng, Weilun Feng +5
Feed-forward 3D reconstruction models, represented by Visual Geometry Grounded Transformer (VGGT), jointly predict multiple visual geometry tasks such as depth estimation, camera p…
Echo-Forcing: A Scene Memory Framework for Interactive Long Video Generation
Mingqiang Wu, Weilun Feng, Zhefeng Zhang +8
Autoregressive video diffusion models enable open-ended generation through local attention and KV caching. However, existing training-free long-video optimization methods mainly fo…