From the 1 of 6 linked papers with an AI index.
6 papers
Von Mises-Fisher Mixture Model with Dynamic Shrinkage for Realistic Test-Time Transduction
Jiazhen Huang, Zhiming Liu, Changhu Wang +3
A range of methods aim to enhance the performance of vision-language models (VLMs) at test time. Among them, transduction has emerged as a promising paradigm due to its strong comp…
Correcting Visual Blur Induced by Attention Distraction to Reduce Hallucinations: Algorithm and Theory
Quanjiang Li, Zhiming Liu, Wei Luo +2
The paper investigates why multimodal large language models hallucinate objects, linking it to an attention distraction effect similar to human visual blur, and introduces AFIP, a…
What Drives Test-Time Adaptation for CLIP? A Controlled Empirical Study from an Update Perspective
Jiazhen Huang, Xiao Chen, Zhiming Liu +3
Vision-Language Models (VLMs) such as CLIP have become a standard backbone for open-vocabulary recognition, yet their zero-shot predictions remain vulnerable to distribution shifts…
Test-Time Distillation for Continual Model Adaptation
Xiao Chen, Jiazhen Huang, Zhiming Liu +4
Deep neural networks often suffer performance degradation upon deployment due to distribution shifts. Continual Test-Time Adaptation (CTTA) aims to address this issue in an unsuper…
Do All Individual Layers Help? An Empirical Study of Task-Interfering Layers in Vision-Language Models
Zhiming Liu, Yujie Wei, Lei Feng +5
Current VLMs have demonstrated capabilities across a wide range of multimodal tasks. Typically, in a pretrained VLM, all layers are engaged by default to make predictions on downst…
Adaptive Disentangled Representation Learning for Incomplete Multi-View Multi-Label Classification
Quanjiang Li, Zhiming Liu, Tianxiang Xu +2
Multi-view multi-label learning frequently suffers from simultaneous feature absence and incomplete annotations, due to challenges in data acquisition and cost-intensive supervisio…