How I failed machine learning in medical imaging -- shortcomings and recommendations
arXiv:2103.10292 · doi:10.1038/s41746-022-00592-y
Abstract
Medical imaging is an important research field with many opportunities for improving patients' health. However, there are a number of challenges that are slowing down the progress of the field as a whole, such optimizing for publication. In this paper we reviewed several problems related to choosing datasets, methods, evaluation metrics, and publication strategies. With a review of literature and our own analysis, we show that at every step, potential biases can creep in. On a positive note, we also see that initiatives to counteract these problems are already being started. Finally we provide a broad range of recommendations on how to further these address problems in the future. For reproducibility, data and code for our analyses are available on \url{https://github.com/GaelVaroquaux/ml_med_imaging_failures}
References in corpus (7)
- Common pitfalls and recommendations for using machine learning to detect and prognosticate for COVID-19 using chest radiographs and CT scans
- Should we really use post-hoc tests based on mean-ranks?
- On the Value of Out-of-Distribution Testing: An Example of Goodhart's Law
- The Problem with Metrics is a Fundamental Problem for AI
- Accounting for Variance in Machine Learning Benchmarks
- The Scientific Method in the Science of Machine Learning
- Evaluating Progress on Machine Learning for Longitudinal Electronic Healthcare Data
Cited by in corpus (26)
- Artificial intelligence in digital pathology: a systematic review and meta-analysis of diagnostic test accuracy
- Large Language Models Streamline Automated Machine Learning for Clinical Studies
- Explainable, Domain-Adaptive, and Federated Artificial Intelligence in Medicine
- On Leakage in Machine Learning Pipelines
- The Impact of ChatGPT and LLMs on Medical Imaging Stakeholders: Perspectives and Use Cases
- Decentralised, Collaborative, and Privacy-preserving Machine Learning for Multi-Hospital Data
- Detecting Shortcuts in Medical Images -- A Case Study in Chest X-rays
- Integration of nested cross-validation, automated hyperparameter optimization, high-performance computing to reduce and quantify the variance of test performance estimation of deep learning models
- Latent Similarity Identifies Important Functional Connections for Phenotype Prediction
- FetMRQC: a robust quality control system for multi-centric fetal brain MRI
- OnDev-LCT: On-Device Lightweight Convolutional Transformers towards federated learning
- Few-shot learning for COVID-19 Chest X-Ray Classification with Imbalanced Data: An Inter vs. Intra Domain Study
- Rethinking model prototyping through the MedMNIST+ dataset collection
- Scaling up self-supervised learning for improved surgical foundation models
- NESTANets: Stable, accurate and efficient neural networks for analysis-sparse inverse problems
- How You Split Matters: Data Leakage and Subject Characteristics Studies in Longitudinal Brain MRI Analysis
- The NCI Imaging Data Commons as a platform for reproducible research in computational pathology
- Launching Insights: A Pilot Study on Leveraging Real-World Observational Data from the Mayo Clinic Platform to Advance Clinical Research
- Data-Driven Volumetric Image Generation from Surface Structures using a Patient-Specific Deep Leaning Model
- Advances in Automated Fetal Brain MRI Segmentation and Biometry: Insights from the FeTA 2024 Challenge
- Automatic Aorta Segmentation with Heavily Augmented, High-Resolution 3-D ResUNet: Contribution to the SEG.A Challenge
- PULASki: Learning inter-rater variability using statistical distances to improve probabilistic segmentation
- Reflections on "Can AI Understand Our Universe?"
- Predicting breast cancer with AI for individual risk-adjusted MRI screening and early detection
- Classical Autoencoder Distillation of Quantum Adversarial Manipulations
- Impacts of Data Splitting Strategies on Parameterized Link Prediction Algorithms