4 papers
6 Fingers, 1 Kidney: Natural Adversarial Medical Images Reveal Critical Weaknesses of Vision-Language Models
Leon Mayer, Piotr Kalinowski, Caroline Ebersbach +6
Vision-language models (VLMs) are increasingly integrated into clinical workflows. However, existing benchmarks primarily assess performance on common anatomical presentations and…
Physics-IQ Verified
Tim Rädsch, Yuki M Asano, Hilde Kuehne +4
Video generative models ( VGMs) have become a new frontier that can be used not just for video generation but for a multitude of downstream tasks, including world modeling. To adva…
Current validation practice undermines surgical AI development
Annika Reinke, Ziying O. Li, Minu D. Tizabi +97
Surgical data science (SDS) is rapidly advancing, yet clinical adoption of artificial intelligence (AI) in surgery remains limited, with inadequate validation as an important contr…
Bridging vision language model (VLM) evaluation gaps with a framework for scalable and cost-effective benchmark generation
Tim Rädsch, Leon Mayer, Simon Pavicic +8
Reliable evaluation of AI models is critical for scientific progress and practical application. While existing VLM benchmarks provide general insights into model capabilities, thei…