9 papers
On the Robustness of Temporal Vision-Language Models for Surgical Endoscopy Videos
Darakshan Rashid, Raza Imam, Ufaq Khan +9
Temporal vision-language models (TVLMs) offer a reusable, prompt-based interface for surgical video understanding, yet, their robustness under clinically realistic acquisition arti…
6 Fingers, 1 Kidney: Natural Adversarial Medical Images Reveal Critical Weaknesses of Vision-Language Models
Leon Mayer, Piotr Kalinowski, Caroline Ebersbach +6
Vision-language models (VLMs) are increasingly integrated into clinical workflows. However, existing benchmarks primarily assess performance on common anatomical presentations and…
Towards Global AI-Driven Cervical Cancer Screening
Thuy Nuong Tran, Ãmer Sümer, Evangelia Christodoulou +15
The global elimination of cervical cancer is a key public health goal set by the World Health Organization (WHO), with screening programs reducing mortality by up to 80%. However,…
Current validation practice undermines surgical AI development
Annika Reinke, Ziying O. Li, Minu D. Tizabi +97
Surgical data science (SDS) is rapidly advancing, yet clinical adoption of artificial intelligence (AI) in surgery remains limited, with inadequate validation as an important contr…
Performance uncertainty in medical image analysis: a large-scale investigation of confidence intervals
Pascaline André, Charles Heitz, Evangelia Christodoulou +10
Performance uncertainty quantification is essential for reliable validation and eventual clinical translation of medical imaging artificial intelligence (AI). Confidence intervals…
False Promises in Medical Imaging AI? Assessing Validity of Outperformance Claims
Evangelia Christodoulou, Annika Reinke, Pascaline Andrè +23
Performance comparisons are fundamental in medical imaging Artificial Intelligence (AI) research, often driving claims of superiority based on relative improvements in common perfo…