20 papers
On the Robustness of Temporal Vision-Language Models for Surgical Endoscopy Videos
Darakshan Rashid, Raza Imam, Ufaq Khan +9
Temporal vision-language models (TVLMs) offer a reusable, prompt-based interface for surgical video understanding, yet, their robustness under clinically realistic acquisition arti…
From Multi-Resolution Cells to Gigapixel Whole Slide Images Foundation Model for Computational Pathology
Basit Alawode, Moshira Ali Abdalla, Dwarikanath Mahapatra +3
Vision Transformers (ViTs) and their hierarchical variants have achieved strong performance in Computational Pathology (CPath). However, most are pre-trained on single-resolution W…
Induce to Empower: Improving Lightweight Baselines via Foundation Model Induction for Generalized Polyp Segmentation
Shivanshu Agnihotri, Snehashis Majhi, Deepak Ranjan Nayak +2
Automated polyp segmentation in colonoscopy continues to pose challenges due to substantial appearance variations and indistinct polyp boundaries. Although emerging foundation mode…
Can Experts Adapt Without Training? On Test-Time Modality Generalization in MVLMs
Raza Imam, Darakshan Rashid, Yutong Xie +3
Medical vision-language models (MVLMs) promise broad zero-shot generalization, yet their reliability collapses when confronted with unseen modalities and domains, precisely where c…
EnTrust: Modeling Inter-Modal Conflict for Trustworthy Multimodal Medical Image Analysis
Dwarikanath Mahapatra, Abhijit Das, Behzad Bozorgtabar +5
Multimodal medical imaging fuses complementary anatomical and functional information, yet modalities frequently disagree in pathologically heterogeneous regions. Current segmentati…
Graph-of-Differences: Anatomy-Structured Difference Alignment for Medical Image Re-Identification
Nichula Wasalathilaka, Abhijit Das, Imran Razzak +1
Medical image re-identification (MedReID) enables longitudinal patient linkage but remains vulnerable to shortcut learning and often produces decisions that clinicians cannot audit…