3 papers
cs.CV2026
MATEX: Multi-scale Attention and Text-guided Explainability of Medical Vision-Language Models
Muhammad Imran, Chi Lee, Yugyung Lee
We introduce MATEX (Multi-scale Attention and Text-guided Explainability), a novel framework that advances interpretability in medical vision-language models by incorporating anato…
cs.CV2026
Predicting When to Trust Vision-Language Models for Spatial Reasoning
Muhammad Imran, Yugyung Lee
Vision-Language Models (VLMs) demonstrate impressive capabilities across multimodal tasks, yet exhibit systematic spatial reasoning failures, achieving only 49% (CLIP) to 54% (BLIP…
cs.CV2025
Multi-Modal Interpretability for Enhanced Localization in Vision-Language Models
Muhammad Imran, Yugyung Lee
Recent advances in vision-language models have significantly expanded the frontiers of automated image analysis. However, applying these models in safety-critical contexts remains…