7 papers
CAR-MIL: Counterfactual Attention Regularization for Multiple Instance Learning
Imane Chraki, Pierre Marza, Stergios Christodoulidis +1
Multiple Instance Learning (MIL) is widely used for weakly supervised learning, particularly in digital pathology, where fine-grained annotations are costly. Most MIL methods aggre…
Medical Context Distorts Decisions in Clinical Vision Language Models
David Restrepo, Ira Ktena, Maria Vakalopoulou +2
Vision-language models (VLMs) are increasingly proposed for clinical decision support, yet their reliability in real-world scenarios that require integrating both visual and textua…
On the Cone Effect and Modality Gap in Medical Vision-Language Embeddings
David Restrepo, Miguel L Martins, Chenwei Wu +5
Vision-Language Models (VLMs) exhibit a characteristic "cone effect" in which nonlinear encoders map embeddings into highly concentrated regions of the representation space, contri…
GATE-AD: Graph Attention Network Encoding For Few-Shot Industrial Visual Anomaly Detection
Aggelos Psiris, Yannis Panagakis, Maria Vakalopoulou +1
Few-Shot Industrial Visual Anomaly Detection (FS-IVAD) comprises a critical task in modern manufacturing settings, where automated product inspection systems need to identify rare…
Information Maximization for Long-Tailed Semi-Supervised Domain Generalization
Leo Fillioux, Omprakash Chakraborty, Quentin Gopée +6
Semi-supervised domain generalization (SSDG) has recently emerged as an appealing alternative to tackle domain generalization when labeled data is scarce but unlabeled samples acro…
CDG-MAE: Cross-view Masked Modeling using Diffusion Generated Views
Varun Belagali, Pierre Marza, Srikar Yellapragada +7
Cross-view masked autoencoding has emerged as a powerful pretext task for learning dense correspondences, which are essential for applications such as video label propagation. The…