11 papers
RADMI: Latent Information Aggregation as a Proxy for Model Uncertainty
William Stevens, Mohit Prabhushankar, Ghassan AlRegib
Epistemic uncertainty estimation is essential for identifying regions where deep learning system outputs may be unreliable. However, existing approaches require computationally exp…
Information Router for Mitigating Modality Dominance in Vision-Language Models
Seulgi Kim, Mohit Prabhushankar, Ghassan AlRegib
Vision Language models (VLMs) have demonstrated strong performance across a wide range of benchmarks, yet they often suffer from modality dominance, where predictions rely dispropo…
Hierarchical and Multimodal Data for Daily Activity Understanding
Ghazal Kaviani, Yavuz Yarici, Seulgi Kim +4
Daily Activity Recordings for Artificial Intelligence (DARai, pronounced "Dahr-ree") is a multimodal, hierarchically annotated dataset constructed to understand human activities in…
BALD-SAM: Disagreement-based Active Prompting in Interactive Segmentation
Prithwijit Chowdhury, Mohit Prabhushankar, Ghassan AlRegib
The Segment Anything Model (SAM) has revolutionized interactive segmentation through spatial prompting. While existing work primarily focuses on automating prompts in various setti…
Gradient based Severity Labeling for Biomarker Classification in OCT
Kiran Kokilepersaud, Mohit Prabhushankar, Ghassan AlRegib +2
In this paper, we propose a novel selection strategy for contrastive learning for medical images. On natural images, contrastive learning uses augmentations to select positive and…
Countering Multi-modal Representation Collapse through Rank-targeted Fusion
Seulgi Kim, Kiran Kokilepersaud, Mohit Prabhushankar +1
Multi-modal fusion methods often suffer from two types of representation collapse: feature collapse where individual dimensions lose their discriminative power (as measured by eige…