3 papers
cs.CV2026
Mitigating Visual Degradation in MLLMs via Spatial-Spectral Visual Anchor Learning
Qianlong Yang, Bowen Ye, Xianda Guo +4
Despite the progress of multimodal large language models (MLLMs), they continue to exhibit deficiencies in visual perception. Following visual instruction tuning, internal MLLM rep…
cs.CV2025
Class-Aware Prototype Learning with Negative Contrast for Test-Time Adaptation of Vision-Language Models
Xiaozhen Qiao, Jingkai Zhao, Yuqiu Jiang +4
Vision-Language Models (VLMs) demonstrate impressive zero-shot generalization through large-scale image-text pretraining, yet their performance can drop once the deployment distrib…
cs.CV2025
Bidirectional Prototype-Reward co-Evolution for Test-Time Adaptation of Vision-Language Models
Xiaozhen Qiao, Peng Huang, Jiakang Yuan +6
Test-time adaptation (TTA) is crucial in maintaining performance of Vision Language Models (VLMs) when facing distribution shifts, particularly when the source data or target label…