6 papers
Efficiency Follows Global-Local Decoupling
Zhenyu Yang, Gensheng Pei, Tao Chen +4
Modern vision models must capture image-level context without sacrificing local detail while remaining computationally affordable. We revisit this tradeoff and advance a simple pri…
Towards Remote Sensing Change Detection with Neural Memory
Zhenyu Yang, Gensheng Pei, Yazhou Yao +3
Remote sensing change detection is essential for environmental monitoring, urban planning, and related applications. However, current methods often struggle to capture long-range d…
AbductiveMLLM: Boosting Visual Abductive Reasoning Within MLLMs
Boyu Chang, Qi Wang, Xi Guo +3
Visual abductive reasoning (VAR) is a challenging task that requires AI systems to infer the most likely explanation for incomplete visual observations. While recent MLLMs develop…
Decouple before Align: Visual Disentanglement Enhances Prompt Tuning
Fei Zhang, Tianfei Zhou, Jiangchao Yao +3
Prompt tuning (PT), as an emerging resource-efficient fine-tuning paradigm, has showcased remarkable effectiveness in improving the task-specific transferability of vision-language…
Seeing What Matters: Empowering CLIP with Patch Generation-to-Selection
Gensheng Pei, Tao Chen, Yujia Wang +4
The CLIP model has demonstrated significant advancements in aligning visual and language modalities through large-scale pre-training on image-text pairs, enabling strong zero-shot…
On-Road Object Importance Estimation: A New Dataset and A Model with Multi-Fold Top-Down Guidance
Zhixiong Nan, Yilong Chen, Tianfei Zhou +1
This paper addresses the problem of on-road object importance estimation, which utilizes video sequences captured from the driver's perspective as the input. Although this problem…