Reading Race: AI Recognises Patient's Racial Identity In Medical Images
arXiv:2107.10356 · doi:10.1016/S2589-7500(22)00063-2
Abstract
Background: In medical imaging, prior studies have demonstrated disparate AI performance by race, yet there is no known correlation for race on medical imaging that would be obvious to the human expert interpreting the images. Methods: Using private and public datasets we evaluate: A) performance quantification of deep learning models to detect race from medical images, including the ability of these models to generalize to external environments and across multiple imaging modalities, B) assessment of possible confounding anatomic and phenotype population features, such as disease distribution and body habitus as predictors of race, and C) investigation into the underlying mechanism by which AI models can recognize race. Findings: Standard deep learning models can be trained to predict race from medical images with high performance across multiple imaging modalities. Our findings hold under external validation conditions, as well as when models are optimized to perform clinically motivated tasks. We demonstrate this detection is not due to trivial proxies or imaging-related surrogate covariates for race, such as underlying disease distribution. Finally, we show that performance persists over all anatomical regions and frequency spectrum of the images suggesting that mitigation efforts will be challenging and demand further study. Interpretation: We emphasize that model ability to predict self-reported race is itself not the issue of importance. However, our findings that AI can trivially predict self-reported race -- even from corrupted, cropped, and noised medical images -- in a setting where clinical experts cannot, creates an enormous risk for all model deployments in medical imaging: if an AI model secretly used its knowledge of self-reported race to misclassify all Black patients, radiologists would not be able to tell using the same data the model has access to.
Cited by in corpus (21)
- Detecting Shortcut Learning for Fair Medical AI using Shortcut Testing
- Risk of Bias in Chest Radiography Deep Learning Foundation Models
- Black-Box Access is Insufficient for Rigorous AI Audits
- Synthetically Enhanced: Unveiling Synthetic Data's Potential in Medical Imaging Research
- Using generative AI to investigate medical imagery models and datasets
- Detecting Shortcuts in Medical Images -- A Case Study in Chest X-rays
- Towards objective and systematic evaluation of bias in artificial intelligence for medical imaging
- Interpretable machine learning for time-to-event prediction in medicine and healthcare
- In the Picture: Medical Imaging Datasets, Artifacts, and their Living Review
- Report of the Medical Image De-Identification (MIDI) Task Group -- Best Practices and Recommendations
- Data Poisoning Vulnerabilities Across Healthcare AI Architectures: A Security Threat Analysis
- Generative AI in Health Economics and Outcomes Research: A Taxonomy of Key Definitions and Emerging Applications, an ISPOR Working Group Report
- Dataset Distribution Impacts Model Fairness: Single vs. Multi-Task Learning
- Using Backbone Foundation Model for Evaluating Fairness in Chest Radiography Without Demographic Data
- Estimating and Controlling for Equalized Odds via Sensitive Attribute Predictors
- Revisiting Hidden Representations in Transfer Learning for Medical Imaging
- The Patient/Industry Trade-off in Medical Artificial Intelligence
- Causal thinking for decision making on Electronic Health Records: why and how
- Augmenting Chest X-ray Datasets with Non-Expert Annotations
- Robustness and sex differences in skin cancer detection: logistic regression vs CNNs
- Fuzzy Accuracy Compensates for Label Subjectivity in Classification of Skin Tone Using Wearable Photoplethysmography Signals