8 papers
nnActive: A Framework for Evaluation of Active Learning in 3D Biomedical Segmentation
Carsten T. Lüth, Jeremias Traub, Kim-Celine Kahl +6
Semantic segmentation is crucial for various biomedical applications, yet its reliance on large annotated datasets presents a bottleneck due to the high cost and specialized expert…
SURE-VQA: Systematic Understanding of Robustness Evaluation in Medical VQA Tasks
Kim-Celine Kahl, Selen Erkan, Jeremias Traub +4
Vision-Language Models (VLMs) have great potential in medical tasks, like Visual Question Answering (VQA), where they could act as interactive assistants for both patients and clin…
RadioActive: 3D Radiological Interactive Segmentation Benchmark
Constantin Ulrich, Tassilo Wald, Emily Tempus +3
Effortless and precise segmentation with minimal clinician effort could greatly streamline clinical workflows. Recent interactive segmentation models, inspired by METAs Segment Any…
Bridging vision language model (VLM) evaluation gaps with a framework for scalable and cost-effective benchmark generation
Tim Rädsch, Leon Mayer, Simon Pavicic +8
Reliable evaluation of AI models is critical for scientific progress and practical application. While existing VLM benchmarks provide general insights into model capabilities, thei…
Application-driven Validation of Posteriors in Inverse Problems
Tim J. Adler, Jan-Hinrich Nölke, Annika Reinke +8
Current deep learning-based solutions for image analysis tasks are commonly incapable of handling problems to which multiple different plausible solutions exist. In response, poste…
Touchstone Benchmark: Are We on the Right Way for Evaluating AI Algorithms for Medical Segmentation?
Pedro R. A. S. Bassi, Wenxuan Li, Yucheng Tang +50
How can we test AI performance? This question seems trivial, but it isn't. Standard benchmarks often have problems such as in-distribution and small-size test sets, oversimplified…