10 papers
Can Experts Adapt Without Training? On Test-Time Modality Generalization in MVLMs
Raza Imam, Darakshan Rashid, Yutong Xie +3
Medical vision-language models (MVLMs) promise broad zero-shot generalization, yet their reliability collapses when confronted with unseen modalities and domains, precisely where c…
CA-GCL: Cross-Anatomy Global-Local Contrastive Learning for Robust 3D Medical Image Understanding
Hanwen Zhang, Yao Liu, Die Dai +4
Fine-grained Vision-Language Pre-training (FVLP) demonstrates significant potential in 3D medical image understanding by aligning anatomy-level visual representations with correspo…
Towards Physically Consistent 4D Scene Reconstruction for Closed-loop Autonomous Driving Simulation
Bowyn Tan, Yutong Xie, Bai Huang +5
High-fidelity street scene reconstruction is pivotal for end-to-end autonomous driving simulation, where novel-view synthesis (NVS) and time-varying information modeling are two fu…
Finding the Correct Visual Evidence Without Forgetting: Mitigating Hallucination in LVLMs via Inter-Layer Visual Attention Discrepancy
Yutong Xie, Zhenglin Hua, Ran Wang +3
Large Vision-Language Models (LVLMs) have shown remarkable performance on a wide range of vision-language tasks. Despite this progress, they are still prone to hallucination, gener…
TriALS: Triphasic-Aided Liver Lesion Segmentation Benchmark in Non-Contrast CT
Marawan Elbatel, Mohamed Ghonim, Jiaji Mao +62
Automated segmentation of liver lesions on non-contrast computed tomography (NCCT) is clinically important but fundamentally challenging, particularly in low-resource settings acro…
Beyond the Global Scores: Fine-Grained Token Grounding as a Robust Detector of LVLM Hallucinations
Tuan Dung Nguyen, Minh Khoi Ho, Qi Chen +8
Large vision-language models (LVLMs) achieve strong performance on visual reasoning tasks but remain highly susceptible to hallucination. Existing detection methods predominantly r…