4 papers · 1 filter
Grounding Isn't Knowing: Do VLMs Need Object Localization for Spatial Reasoning?
Xiwei Liu, Yulong Li, Xinlin Zhuang +5
Vision-language models (VLMs) can answer spatial questions, yet the mechanisms connecting object grounding to spatial reasoning remain poorly understood. It is underexplored whethe…
ClinCoT: Clinical-Aware Visual Chain-of-Thought for Medical Vision Language Models
Xiwei Liu, Yulong Li, Xinlin Zhuang +5
Medical Vision-Language Models have shown promising potential in clinical decision support, yet they remain prone to factual hallucinations due to insufficient grounding in localiz…
Towards Robust Visual Continual Learning with Multi-Prototype Supervision
Xiwei Liu, Yulong Li, Yichen Li +4
Language-guided supervision, which utilizes a frozen semantic target from a Pretrained Language Model (PLM), has emerged as a promising paradigm for visual Continual Learning (CL).…
SSS: Semi-Supervised SAM-2 with Efficient Prompting for Medical Imaging Segmentation
Hongjie Zhu, Xiwei Liu, Rundong Xue +5
In the era of information explosion, efficiently leveraging large-scale unlabeled data while minimizing the reliance on high-quality pixel-level annotations remains a critical chal…