5 citations · 5 across the 12 of their papers we have counts for
4 papers · 1 filter
Not Another Text Benchmark: Putting the "Visual" Back in Visual Question Answering for Large Video Models
Rwiddhi Chakraborty, Yinong, Wang +7
Large video models have exhibited impressive performance on a wide range of visual question answering tasks, owing to the rise of powerful, pretrained text and vision encoders. The…
SMILE: Self-Explainable Multimodal Information Bottleneck for Medical Diagnosis
Yuqing Yang, Alexander Schmatz, Zhaozhao Ma +3
Explainability is increasingly seen as a crucial requirement in AI-based medical diagnosis, particularly in safety-critical clinical decision-making. Most existing explainability m…
Random Window Augmentations for Deep Learning Robustness in CT and Liver Tumor Segmentation
Eirik A. Østmo, Kristoffer K. Wickstrøm, Keyur Radiya +3
Contrast-enhanced Computed Tomography (CT) is important for diagnosis and treatment planning for various medical conditions. Deep learning (DL) based segmentation models may enable…
From Colors to Classes: Emergence of Concepts in Vision Transformers
Teresa Dorszewski, Lenka Tětková, Robert Jenssen +2
Vision Transformers (ViTs) are increasingly utilized in various computer vision tasks due to their powerful representation capabilities. However, it remains understudied how ViTs p…