Stop Explaining Black Box Machine Learning Models for High Stakes Decisions and Use Interpretable Models Instead
arXiv:1811.10154
Abstract
Black box machine learning models are currently being used for high stakes decision-making throughout society, causing problems throughout healthcare, criminal justice, and in other domains. People have hoped that creating methods for explaining these black box models will alleviate some of these problems, but trying to \textit{explain} black box models, rather than creating models that are \textit{interpretable} in the first place, is likely to perpetuate bad practices and can potentially cause catastrophic harm to society. There is a way forward -- it is to design models that are inherently interpretable. This manuscript clarifies the chasm between explaining black boxes and using inherently interpretable models, outlines several key reasons why explainable black boxes should be avoided in high-stakes decisions, identifies challenges to interpretable machine learning, and provides several example applications where interpretable models could potentially replace black box models in criminal justice, healthcare, and computer vision.
Author's pre-publication version of a 2019 Nature Machine Intelligence article. Shorter Version was published in NIPS 2018 Workshop on Critiquing and Correcting Trends in Machine Learning. Expands also on NSF Statistics at a Crossroads Webinar
References in corpus (8)
- European Union regulations on algorithmic decision-making and a "right to explanation"
- Confounding variables can degrade generalization performance of radiological deep learning models
- Classifier Technology and the Illusion of Progress
- Deep Learning for Case-Based Reasoning through Prototypes: A Neural Network that Explains Its Predictions
- An Interpretable Model with Globally Consistent Explanations for Credit Risk
- Boolean Decision Rules via Column Generation
- Random Forests, Decision Trees, and Categorical Predictors: The "Absent Levels" Problem
- Interpretable Two-level Boolean Rule Learning for Classification
Cited by in corpus (22)
- Model-Agnostic Counterfactual Explanations for Consequential Decisions
- Explainable AI: current status and future directions
- Quantifying Model Complexity via Functional Decomposition for Better Post-Hoc Interpretability
- AutoAIViz: Opening the Blackbox of Automated Artificial Intelligence with Conditional Parallel Coordinates
- Checklist for responsible deep learning modeling of medical images based on COVID-19 detection studies
- In AI We Trust? Factors That Influence Trustworthiness of AI-infused Decision-Making Processes
- On the Art and Science of Machine Learning Explanations
- Enforcing Interpretability and its Statistical Impacts: Trade-offs between Accuracy and Interpretability
- Explaining Deep Classification of Time-Series Data with Learned Prototypes
- AutoScore-Imbalance: An interpretable machine learning tool for development of clinical scores with rare events data
- Proposed Guidelines for the Responsible Use of Explainable Machine Learning
- Hybrid Predictive Model: When an Interpretable Model Collaborates with a Black-box Model
- Region Comparison Network for Interpretable Few-shot Image Classification
- Explaining Deep Neural Networks using Unsupervised Clustering
- A Framework for Democratizing AI
- Learning Fair Rule Lists
- Interpreting Neural Networks Using Flip Points
- An Empirical Study of Explainable AI Techniques on Deep Learning Models For Time Series Tasks
- A framework for predicting, interpreting, and improving Learning Outcomes
- Fair and Responsible AI: A Focus on the Ability to Contest
- A Multi-Objective Anytime Rule Mining System to Ease Iterative Feedback from Domain Experts
- Distilling Interpretable Models into Human-Readable Code