4 papers
ERUnderstand: Evaluating Vision-Language Models on Structured ER Diagrams
Ali Ansari, Yasmin Mohammadi, Farnoush Nili +3
Entity-Relationship Diagrams (ERDs) are central to conceptual database design, yet they are typically available only as rendered images rather than machine-readable schemas, limiti…
Validation of Whole-Slide Foundation Models for Image Retrieval in TCGA Data
Tianhao Lei, Parsa Esmaeilkhani, Saghir Alfasly +5
Foundation models are reshaping computational histopathology, yet their value for whole-slide image retrieval relative to strong patch-based and supervised aggregation baselines re…
Logit Lens Supervision for Patch-Level Explanations in Vision-Language Models
Parsa Esmaeilkhani, Longin Jan Latecki
Modern autoregressive Vision-Language Models (VLMs) can generate fluent answers while their visual-token representations become weakly tied to the image regions from which they ori…
Direct Visual Grounding by Directing Attention of Visual Tokens
Parsa Esmaeilkhani, Longin Jan Latecki
Vision Language Models (VLMs) mix visual tokens and text tokens. A puzzling issue is the fact that visual tokens most related to the query receive little to no attention in the fin…