DC-Check: A Data-Centric AI checklist to guide the development of reliable machine learning systems
arXiv:2211.05764 · doi:10.1109/TAI.2023.3345805
Abstract
While there have been a number of remarkable breakthroughs in machine learning (ML), much of the focus has been placed on model development. However, to truly realize the potential of machine learning in real-world settings, additional aspects must be considered across the ML pipeline. Data-centric AI is emerging as a unifying paradigm that could enable such reliable end-to-end pipelines. However, this remains a nascent area with no standardized framework to guide practitioners to the necessary data-centric considerations or to communicate the design of data-centric driven ML systems. To address this gap, we propose DC-Check, an actionable checklist-style framework to elicit data-centric considerations at different stages of the ML pipeline: Data, Training, Testing, and Deployment. This data-centric lens on development aims to promote thoughtfulness and transparency prior to system development. Additionally, we highlight specific data-centric AI challenges and research opportunities. DC-Check is aimed at both practitioners and researchers to guide day-to-day development. As such, to easily engage with and use DC-Check and associated resources, we provide a DC-Check companion website (https://www.vanderschaar-lab.com/dc-check/). The website will also serve as an updated resource as methods and tooling evolve over time.
Main paper: 11 pages, supplementary & case studies follow
References in corpus (12)
- Common pitfalls and recommendations for using machine learning to detect and prognosticate for COVID-19 using chest radiographs and CT scans
- A Model of Inductive Bias Learning
- Probabilistic Backpropagation for Scalable Learning of Bayesian Neural Networks
- Training Convolutional Networks with Noisy Labels
- An introduction to domain adaptation and transfer learning
- Degenerate Feedback Loops in Recommender Systems
- Checklist for responsible deep learning modeling of medical images based on COVID-19 detection studies
- AutoPrognosis 2.0: Democratizing Diagnostic and Prognostic Modeling in Healthcare with Automated Machine Learning
- A Data Quality-Driven View of MLOps
- AlphaClean: Automatic Generation of Data Cleaning Pipelines
- DECAF: Generating Fair Synthetic Data Using Causally-Aware Generative Networks
- Synthcity: facilitating innovative use cases of synthetic data in different data modalities