Can we trust deep learning models diagnosis? The impact of domain shift in chest radiograph classification
arXiv:1909.01940 · doi:10.1007/978-3-030-62469-9_7
Abstract
While deep learning models become more widespread, their ability to handle unseen data and generalize for any scenario is yet to be challenged. In medical imaging, there is a high heterogeneity of distributions among images based on the equipment that generates them and their parametrization. This heterogeneity triggers a common issue in machine learning called domain shift, which represents the difference between the training data distribution and the distribution of where a model is employed. A high domain shift tends to implicate in a poor generalization performance from the models. In this work, we evaluate the extent of domain shift on four of the largest datasets of chest radiographs. We show how training and testing with different datasets (e.g., training in ChestX-ray14 and testing in CheXpert) drastically affects model performance, posing a big question over the reliability of deep learning models trained on public datasets. We also show that models trained on CheXpert and MIMIC-CXR generalize better to other datasets.
10 pages, 3 figures
References in corpus (1)
Cited by in corpus (13)
- Domain Adaptation for Medical Image Analysis: A Survey
- How I failed machine learning in medical imaging -- shortcomings and recommendations
- The reliability of a deep learning model in clinical out-of-distribution MRI data: a multicohort study
- Normalization of breast MRIs using Cycle-Consistent Generative Adversarial Networks
- Detecting Shortcuts in Medical Images -- A Case Study in Chest X-rays
- Mind the Gap: Federated Learning Broadens Domain Generalization in Diagnostic AI Models
- Preserving privacy in domain transfer of medical AI models comes at no performance costs: The integral role of differential privacy
- Augmentation-based Domain Generalization and Joint Training from Multiple Source Domains for Whole Heart Segmentation
- LA-CaRe-CNN: Cascading Refinement CNN for Left Atrial Scar Segmentation
- Space-scale Exploration of the Poor Reliability of Deep Learning Models: the Case of the Remote Sensing of Rooftop Photovoltaic Systems
- Federated Deep AUC Maximization for Heterogeneous Data with a Constant Communication Complexity
- MIMM-X: Disentangling Spurious Correlations for Medical Image Analysis
- Requirement analysis for an artificial intelligence model for the diagnosis of the COVID-19 from chest X-ray data