most citedTouchstone Benchmark: Are We on the Right Way for Evaluating AI Algorithms for Medical Segmentation?

2 citations · 3 across the 6 of their papers we have counts for

collaborators

6 papers

eess.IV2025

Promptable Longitudinal Lesion Segmentation in Whole-Body CT

Yannick Kirchhoff, Maximilian Rokuss, Fabian Isensee +1

Accurate segmentation of lesions in longitudinal whole-body CT is essential for monitoring disease progression and treatment response. While automated methods benefit from incorpor…

cs.CV2025

A Multi-Stage Fine-Tuning and Ensembling Strategy for Pancreatic Tumor Segmentation in Diagnostic and Therapeutic MRI

Omer Faruk Durugol, Maximilian Rokuss, Yannick Kirchhoff +1

Automated segmentation of Pancreatic Ductal Adenocarcinoma (PDAC) from MRI is critical for clinical workflows but is hindered by poor tumor-tissue contrast and a scarcity of annota…

cs.CV2025

Towards Interactive Lesion Segmentation in Whole-Body PET/CT with Promptable Models

Maximilian Rokuss, Yannick Kirchhoff, Fabian Isensee +1

Whole-body PET/CT is a cornerstone of oncological imaging, yet accurate lesion segmentation remains challenging due to tracer heterogeneity, physiological uptake, and multi-center…

cs.CV2025

Temporal Flow Matching for Learning Spatio-Temporal Trajectories in 4D Longitudinal Medical Imaging

Nico Albert Disch, Yannick Kirchhoff, Robin Peretzke +5

Understanding temporal dynamics in medical imaging is crucial for applications such as disease progression modeling, treatment planning and anatomical development tracking. However…

cs.CV20241 cited

Scaling nnU-Net for CBCT Segmentation

Fabian Isensee, Yannick Kirchhoff, Lars Kraemer +3

This paper presents our approach to scaling the nnU-Net framework for multi-structure segmentation on Cone Beam Computed Tomography (CBCT) images, specifically in the scope of the…

cs.CV20242 cited

Touchstone Benchmark: Are We on the Right Way for Evaluating AI Algorithms for Medical Segmentation?

Pedro R. A. S. Bassi, Wenxuan Li, Yucheng Tang +50

How can we test AI performance? This question seems trivial, but it isn't. Standard benchmarks often have problems such as in-distribution and small-size test sets, oversimplified…