most citedCurrent validation practice undermines surgical AI development

1 citations · 1 across the 2 of their papers we have counts for

collaborators

6 papers

cs.CV2026

6 Fingers, 1 Kidney: Natural Adversarial Medical Images Reveal Critical Weaknesses of Vision-Language Models

Leon Mayer, Piotr Kalinowski, Caroline Ebersbach +6

Vision-language models (VLMs) are increasingly integrated into clinical workflows. However, existing benchmarks primarily assess performance on common anatomical presentations and…

q-bio.OT20261 cited

Current validation practice undermines surgical AI development

Annika Reinke, Ziying O. Li, Minu D. Tizabi +97

Surgical data science (SDS) is rapidly advancing, yet clinical adoption of artificial intelligence (AI) in surgery remains limited, with inadequate validation as an important contr…

cs.CV2025

False Promises in Medical Imaging AI? Assessing Validity of Outperformance Claims

Evangelia Christodoulou, Annika Reinke, Pascaline Andrè +23

Performance comparisons are fundamental in medical imaging Artificial Intelligence (AI) research, often driving claims of superiority based on relative improvements in common perfo…

cs.CV2025

Challenging Vision-Language Models with Surgical Data: A New Dataset and Broad Benchmarking Study

Leon Mayer, Tim Rädsch, Dominik Michael +8

While traditional computer vision models have historically struggled to generalize to endoscopic domains, the emergence of foundation models has shown promising cross-domain perfor…

cs.CV2025

Large-scale Self-supervised Video Foundation Model for Intelligent Surgery

Shu Yang, Fengtao Zhou, Leon Mayer +16

Computer-Assisted Intervention (CAI) has the potential to revolutionize modern surgery, with surgical scene understanding serving as a critical component in supporting decision-mak…

cs.CV2025

Bridging vision language model (VLM) evaluation gaps with a framework for scalable and cost-effective benchmark generation

Tim Rädsch, Leon Mayer, Simon Pavicic +8

Reliable evaluation of AI models is critical for scientific progress and practical application. While existing VLM benchmarks provide general insights into model capabilities, thei…