5 papers
Challenging Vision-Language Models with Surgical Data: A New Dataset and Broad Benchmarking Study
Leon Mayer, Tim Rädsch, Dominik Michael +8
While traditional computer vision models have historically struggled to generalize to endoscopic domains, the emergence of foundation models has shown promising cross-domain perfor…
False Promises in Medical Imaging AI? Assessing Validity of Outperformance Claims
Evangelia Christodoulou, Annika Reinke, Pascaline Andrè +23
Performance comparisons are fundamental in medical imaging Artificial Intelligence (AI) research, often driving claims of superiority based on relative improvements in common perfo…
Bridging vision language model (VLM) evaluation gaps with a framework for scalable and cost-effective benchmark generation
Tim Rädsch, Leon Mayer, Simon Pavicic +8
Reliable evaluation of AI models is critical for scientific progress and practical application. While existing VLM benchmarks provide general insights into model capabilities, thei…
SURE-VQA: Systematic Understanding of Robustness Evaluation in Medical VQA Tasks
Kim-Celine Kahl, Selen Erkan, Jeremias Traub +4
Vision-Language Models (VLMs) have great potential in medical tasks, like Visual Question Answering (VQA), where they could act as interactive assistants for both patients and clin…
Confidence intervals uncovered: Are we ready for real-world medical imaging AI?
Evangelia Christodoulou, Annika Reinke, Rola Houhou +19
Medical imaging is spearheading the AI transformation of healthcare. Performance reporting is key to determine which methods should be translated into clinical practice. Frequently…