73 citations · 186 across the 72 of their papers we have counts for
40 papers · 1 filter
VisLens: Single-Pass Interpretable Visual Search for Multimodal LLMs
Jingyi He, Sanghwan Kim, Zeynep Akata
Multimodal large language models (MLLMs) struggle with fine-grained Visual Search, the task of locating small or rare objects in high-resolution images. Existing remedies fall into…
From Accuracy to Visual Dependence: Auditing and Filtering Modality Collapse in Traffic VideoQA
Sena Korkut, María Alejandra Bravo Sarmiento, Sanghwan Kim +1
High benchmark accuracy does not guarantee genuine use of visual evidence. We study this problem in traffic accident Video Question Answering (VideoQA), where correct answers shoul…
UNBOX: Unveiling Black-box visual models with Natural-language
Simone Carnemolla, Chiara Russo, Simone Palazzo +5
Ensuring trustworthiness in open-world visual recognition requires models that are interpretable, fair, and robust to distribution shifts. Yet modern vision systems are increasingl…
Explaining CLIP Zero-shot Predictions Through Concepts
Onat Ozdemir, Anders Christensen, Stephan Alaniz +2
Large-scale vision-language models such as CLIP have achieved remarkable success in zero-shot image recognition, yet their predictions remain largely opaque to human understanding.…
FINER: MLLMs Hallucinate under Fine-grained Negative Queries
Rui Xiao, Sanghwan Kim, Yongqin Xian +2
Multimodal large language models (MLLMs) struggle with hallucinations, particularly with fine-grained queries, a challenge underrepresented by existing benchmarks that focus on coa…
From Drop-off to Recovery: A Mechanistic Analysis of Segmentation in MLLMs
Boyong Wu, Sanghwan Kim, Zeynep Akata
Multimodal Large Language Models (MLLMs) are increasingly applied to pixel-level vision tasks, yet their intrinsic capacity for spatial understanding remains poorly understood. We…