Evaluating AI systems under uncertain ground truth: a case study in dermatology
arXiv:2307.02191 · doi:10.1016/j.media.2025.103556
Abstract
For safety, medical AI systems undergo thorough evaluations before deployment, validating their predictions against a ground truth which is assumed to be fixed and certain. However, this ground truth is often curated in the form of differential diagnoses. While a single differential diagnosis reflects the uncertainty in one expert assessment, multiple experts introduce another layer of uncertainty through disagreement. Both forms of uncertainty are ignored in standard evaluation which aggregates these differential diagnoses to a single label. In this paper, we show that ignoring uncertainty leads to overly optimistic estimates of model performance, therefore underestimating risk associated with particular diagnostic decisions. To this end, we propose a statistical aggregation approach, where we infer a distribution on probabilities of underlying medical condition candidates themselves, based on observed annotations. This formulation naturally accounts for the potential disagreements between different experts, as well as uncertainty stemming from individual differential diagnoses, capturing the entire ground truth uncertainty. Our approach boils down to generating multiple samples of medical condition probabilities, then evaluating and averaging performance metrics based on these sampled probabilities. In skin condition classification, we find that a large portion of the dataset exhibits significant ground truth uncertainty and standard evaluation severely over-estimates performance without providing uncertainty estimates. In contrast, our framework provides uncertainty estimates on common metrics of interest such as top-k accuracy and average overlap, showing that performance can change multiple percentage points. We conclude that, while assuming a crisp ground truth can be acceptable for many AI applications, a more nuanced evaluation protocol should be utilized in medical diagnosis.
References in corpus (15)
- A deep learning system for differential diagnosis of skin diseases
- Deep Label Distribution Learning with Label Ambiguity
- Why rankings of biomedical image analysis competitions should be interpreted with care
- Deep Learning and Glaucoma Specialists: The Relative Importance of Optic Disc Features to Predict Glaucoma Referral in Fundus Photos
- Jury Learning: Integrating Dissenting Voices into Machine Learning Models
- How To Grade a Test Without Knowing the Answers --- A Bayesian Graphical Model for Adaptive Crowdsourcing and Aptitude Testing
- Fine-tuning language models to find agreement among humans with diverse preferences
- Understanding and Utilizing Deep Neural Networks Trained with Noisy Labels
- Disentangling Human Error from the Ground Truth in Segmentation of Medical Images
- Direct Uncertainty Prediction for Medical Second Opinions
- Robust and Efficient Medical Imaging with Self-Supervision
- Measuring Annotator Agreement Generally across Complex Structured, Multi-object, and Free-text Annotation Tasks
- Consistency is Key: Disentangling Label Variation in Natural Language Processing with Intra-Annotator Agreement
- Learning from Label Proportions by Learning with Label Noise
- Robustness to Label Noise Depends on the Shape of the Noise Distribution in Feature Space