Mask of truth: model sensitivity to unexpected regions of medical images
arXiv:2412.04030 · doi:10.1007/s10278-025-01531-5
Abstract
The development of larger models for medical image analysis has led to increased performance. However, it also affected our ability to explain and validate model decisions. Models can use non-relevant parts of images, also called spurious correlations or shortcuts, to obtain high performance on benchmark datasets but fail in real-world scenarios. In this work, we challenge the capacity of convolutional neural networks (CNN) to classify chest X-rays and eye fundus images while masking out clinically relevant parts of the image. We show that all models trained on the PadChest dataset, irrespective of the masking strategy, are able to obtain an Area Under the Curve (AUC) above random. Moreover, the models trained on full images obtain good performance on images without the region of interest (ROI), even superior to the one obtained on images only containing the ROI. We also reveal a possible spurious correlation in the Chaksu dataset while the performances are more aligned with the expectation of an unbiased model. We go beyond the performance analysis with the usage of the explainability method SHAP and the analysis of embeddings. We asked a radiology resident to interpret chest X-rays under different masking to complement our findings with clinical knowledge. Our code is available at https://github.com/TheoSourget/MMC_Masking and https://github.com/TheoSourget/MMC_Masking_EyeFundus
Updated after publication in the Journal of Imaging Informatics in Medicine
References in corpus (16)
- Shortcut Learning in Deep Neural Networks
- Metrics reloaded: Recommendations for image analysis validation
- Deep learning on fundus images detects glaucoma beyond the optic disc
- Impossibility Theorems for Feature Attribution
- CheXmask: a large-scale dataset of anatomical segmentation masks for multi-center chest x-ray images
- Classification of COPD with Multiple Instance Learning
- Detecting Shortcuts in Medical Images -- A Case Study in Chest X-rays
- Investigating the Quality of DermaMNIST and Fitzpatrick17k Dermatological Image Datasets
- LT-ViT: A Vision Transformer for multi-label Chest X-ray classification
- Towards Generalist Foundation Model for Radiology by Leveraging Web-scale 2D&3D Medical Data
- RadEdit: stress-testing biomedical vision models via diffusion image editing
- Counterfactual contrastive learning: robust representations via causal image synthesis
- Navigating the landscape of multimodal AI in medicine: a scoping review on technical challenges and clinical applications
- Model-based Cleaning of the QUILT-1M Pathology Dataset for Text-Conditional Image Synthesis
- Navigating the Maze of Explainable AI: A Systematic Approach to Evaluating Methods and Metrics
- VisionUnite: A Vision-Language Foundation Model for Ophthalmology Enhanced with Clinical Knowledge