Why rankings of biomedical image analysis competitions should be interpreted with care
arXiv:1806.02051 · doi:10.1038/s41467-018-07619-7
Abstract
International challenges have become the standard for validation of biomedical image analysis methods. Given their scientific impact, it is surprising that a critical analysis of common practices related to the organization of challenges has not yet been performed. In this paper, we present a comprehensive analysis of biomedical image analysis challenges conducted up to now. We demonstrate the importance of challenges and show that the lack of quality control has critical consequences. First, reproducibility and interpretation of the results is often hampered as only a fraction of relevant information is typically provided. Second, the rank of an algorithm is generally not robust to a number of variables such as the test data used for validation, the ranking scheme applied and the observers that make the reference annotations. To overcome these problems, we recommend best practice guidelines and define open research questions to be addressed in the future.
Article published in Nature Communications: https://rdcu.be/bRmNr
References in corpus (3)
- Validation, comparison, and combination of algorithms for automatic detection of pulmonary nodules in computed tomography images: the LUNA16 challenge
- A Novel Performance Evaluation Methodology for Single-Target Trackers
- An Empirical Study into Annotator Agreement, Ground Truth Estimation, and Algorithm Evaluation
Cited by in corpus (44)
- Automated Design of Deep Learning Methods for Biomedical Image Segmentation
- The Medical Segmentation Decathlon
- REFUGE Challenge: A Unified Framework for Evaluating Automated Methods for Glaucoma Assessment from Fundus Photographs
- CHAOS Challenge -- Combined (CT-MR) Healthy Abdominal Organ Segmentation
- A large annotated medical image dataset for the development and evaluation of segmentation algorithms
- How I failed machine learning in medical imaging -- shortcomings and recommendations
- VerSe: A Vertebrae Labelling and Segmentation Benchmark for Multi-detector CT Images
- Causality matters in medical imaging
- Medical Deep Learning -- A systematic Meta-Review
- Understanding metric-related pitfalls in image analysis validation
- MedPerf: Open Benchmarking Platform for Medical Artificial Intelligence using Federated Evaluation
- PanNuke Dataset Extension, Insights and Baselines
- The Multi-modality Cell Segmentation Challenge: Towards Universal Solutions
- CrossMoDA 2021 challenge: Benchmark of Cross-Modality Domain Adaptation techniques for Vestibular Schwannoma and Cochlea Segmentation
- Robust deep learning-based semantic organ segmentation in hyperspectral images
- Common Limitations of Image Processing Metrics: A Picture Story
- Cats or CAT scans: transfer learning from natural or medical image source datasets?
- Self-Supervised Pre-Training with Contrastive and Masked Autoencoder Methods for Dealing with Small Datasets in Deep Learning for Medical Imaging
- Robust Medical Instrument Segmentation Challenge 2019
- Uncertainty-aware performance assessment of optical imaging modalities with invertible neural networks
- APIS: A paired CT-MRI dataset for ischemic stroke segmentation challenge
- A Robust Ensemble Algorithm for Ischemic Stroke Lesion Segmentation: Generalizability and Clinical Utility Beyond the ISLES Challenge
- Multi-structure bone segmentation in pediatric MR images with combined regularization from shape priors and adversarial network
- The Federated Tumor Segmentation (FeTS) Challenge
- Report of the Medical Image De-Identification (MIDI) Task Group -- Best Practices and Recommendations
- AGE Challenge: Angle Closure Glaucoma Evaluation in Anterior Segment Optical Coherence Tomography
- The NCI Imaging Data Commons as a platform for reproducible research in computational pathology
- Less is More: Selective Reduction of CT Data for Self-Supervised Pre-Training of Deep Learning Models with Contrastive Learning Improves Downstream Classification Performance
- Deep Learning Based Cardiac MRI Segmentation: Do We Need Experts?
- From Nano to Macro: Overview of the IEEE Bio Image and Signal Processing Technical Committee
- Foundations of the Theory of Performance-Based Ranking
- How can we learn (more) from challenges? A statistical approach to driving future algorithm development
- Architecture Analysis and Benchmarking of 3D U-shaped Deep Learning Models for Thoracic Anatomical Segmentation
- Self-Supervised Learning from Unlabeled Fundus Photographs Improves Segmentation of the Retina
- Evaluating AI systems under uncertain ground truth: a case study in dermatology
- A Framework for Challenge Design: Insight and Deployment Challenges to Address Medical Image Analysis Problems
- Why is the winner the best?
- Instance Segmentation XXL-CT Challenge of a Historic Airplane
- Inflation of test accuracy due to data leakage in deep learning-based classification of OCT images
- Deep learning for image segmentation: veritable or overhyped?
- The Missing Piece: A Case for Pre-Training in 3D Medical Object Detection
- Knee menisci segmentation and relaxometry of 3D ultrashort echo time (UTE) cones MR imaging using attention U-Net with transfer learning
- MixMicrobleedNet: segmentation of cerebral microbleeds using nnU-Net
- mvHOTA: A multi-view higher order tracking accuracy metric to measure spatial and temporal associations in multi-point detection