4 papers
LASER: A Corrective Lens for LVLMs via Visual Attention Preservation and Sink Suppression
Bowen Yuan, Zijian Wang, Yadan Luo +2
Large vision-language models (LVLMs) exhibit strong reasoning ability but suffer from visual forgetting during long-horizon decoding, where attention progressively drifts away from…
GeoExplain: Multimodal Reasoning based on Hierarchy of Visual Information in Street View
Fenghua Cheng, Jinxiang Wang, Sen Wang +2
Multimodal reasoning is a process of understanding, integrating and inferring information across different data modalities. It has recently attracted surging academic attention. Al…
SCORE: Soft Label Compression-Centric Dataset Condensation via Coding Rate Optimization
Bowen Yuan, Yuxia Fu, Zijian Wang +2
Dataset Condensation (DC) aims to obtain a condensed dataset that allows models trained on the condensed dataset to achieve performance comparable to those trained on the full data…
Is Less More? Exploring Token Condensation as Training-free Test-time Adaptation
Zixin Wang, Dong Gong, Sen Wang +2
Contrastive Language-Image Pretraining (CLIP) excels at learning generalizable image representations but often falls short in zero-shot inference on certain downstream datasets. Te…