In the Picture: Medical Imaging Datasets, Artifacts, and their Living Review
arXiv:2501.10727 · doi:10.1145/3715275.3732035
Abstract
Datasets play a critical role in medical imaging research, yet issues such as label quality, shortcuts, and metadata are often overlooked. This lack of attention may harm the generalizability of algorithms and, consequently, negatively impact patient outcomes. While existing medical imaging literature reviews mostly focus on machine learning (ML) methods, with only a few focusing on datasets for specific applications, these reviews remain static -- they are published once and not updated thereafter. This fails to account for emerging evidence, such as biases, shortcuts, and additional annotations that other researchers may contribute after the dataset is published. We refer to these newly discovered findings of datasets as research artifacts. To address this gap, we propose a living review that continuously tracks public datasets and their associated research artifacts across multiple medical imaging applications. Our approach includes a framework for the living review to monitor data documentation artifacts, and an SQL database to visualize the citation relationships between research artifact and dataset. Lastly, we discuss key considerations for creating medical imaging datasets, review best practices for data annotation, discuss the significance of shortcuts and demographic diversity, and emphasize the importance of managing datasets throughout their entire lifecycle. Our demo is publicly available at http://inthepicture.itu.dk/.
ACM Conference on Fairness, Accountability, and Transparency - FAccT 2025
References in corpus (17)
- The Future of Digital Health with Federated Learning
- Domain Adaptation for Medical Image Analysis: A Survey
- Reading Race: AI Recognises Patient's Racial Identity In Medical Images
- Ethical Machine Learning in Health Care
- Quality Control in Crowdsourcing: A Survey of Quality Attributes, Assessment Techniques and Assurance Actions
- ISLES 2022: A multi-center magnetic resonance imaging stroke lesion segmentation dataset
- A Survey on Deep Learning for Skin Lesion Segmentation
- Do Datasets Have Politics? Disciplinary Values in Computer Vision Dataset Development
- Active label cleaning for improved dataset quality under resource constraints
- Using generative AI to investigate medical imagery models and datasets
- Long-Tailed Classification of Thorax Diseases on Chest X-Ray: A New Benchmark Study
- CheXmask: a large-scale dataset of anatomical segmentation masks for multi-center chest x-ray images
- The Rio Hortega University Hospital Glioblastoma dataset: a comprehensive collection of preoperative, early postoperative and recurrence MRI scans (RHUH-GBM)
- Detecting Shortcuts in Medical Images -- A Case Study in Chest X-rays
- Ground Truth Or Dare: Factors Affecting The Creation Of Medical Datasets For Training AI
- Croissant: A Metadata Format for ML-Ready Datasets
- How Does Pruning Impact Long-Tailed Multi-Label Medical Image Classifiers?