The Hidden Assumptions Behind Counterfactual Explanations and Principal Reasons
arXiv:1912.04930 · doi:10.1145/3351095.3372830
Abstract
Counterfactual explanations are gaining prominence within technical, legal, and business circles as a way to explain the decisions of a machine learning model. These explanations share a trait with the long-established "principal reason" explanations required by U.S. credit laws: they both explain a decision by highlighting a set of features deemed most relevant--and withholding others. These "feature-highlighting explanations" have several desirable properties: They place no constraints on model complexity, do not require model disclosure, detail what needed to be different to achieve a different decision, and seem to automate compliance with the law. But they are far more complex and subjective than they appear. In this paper, we demonstrate that the utility of feature-highlighting explanations relies on a number of easily overlooked assumptions: that the recommended change in feature values clearly maps to real-world actions, that features can be made commensurate by looking only at the distribution of the training data, that features are only relevant to the decision at hand, and that the underlying model is stable over time, monotonic, and limited to binary outcomes. We then explore several consequences of acknowledging and attempting to address these assumptions, including a paradox in the way that feature-highlighting explanations aim to respect autonomy, the unchecked power that feature-highlighting explanations grant decision makers, and a tension between making these explanations useful and the need to keep the model hidden. While new research suggests several ways that feature-highlighting explanations can work around some of the problems that we identify, the disconnect between features in the model and actions in the real world--and the subjective choices necessary to compensate for this--must be understood before these techniques can be usefully implemented.
Cited by in corpus (52)
- The Fallacy of AI Functionality
- Preserving Causal Constraints in Counterfactual Explanations for Machine Learning Classifiers
- Towards Unifying Feature Attribution and Counterfactual Explanations: Different Means to the Same End
- Algorithmic recourse under imperfect causal knowledge: a probabilistic approach
- Algorithmic Recourse: from Counterfactual Explanations to Interventions
- Explainable Machine Learning for Public Policy: Use Cases, Gaps, and Research Directions
- Counterfactual Shapley Additive Explanations
- Explaining Data-Driven Decisions made by AI Systems: The Counterfactual Approach
- The Conflict Between Explainable and Accountable Decision-Making Algorithms
- A Hierarchy of Limitations in Machine Learning
- A survey of algorithmic recourse: definitions, formulations, solutions, and prospects
- Amazon SageMaker Clarify: Machine Learning Bias Detection and Explainability in the Cloud
- Towards Robust and Reliable Algorithmic Recourse
- Epistemic values in feature importance methods: Lessons from feminist epistemology
- What-is and How-to for Fairness in Machine Learning: A Survey, Reflection, and Perspective
- Sparse Visual Counterfactual Explanations in Image Space
- Beyond Individualized Recourse: Interpretable and Interactive Summaries of Actionable Recourses
- On Counterfactual Explanations under Predictive Multiplicity
- GAM Coach: Towards Interactive and User-centered Algorithmic Recourse
- Counterfactuals and Causability in Explainable Artificial Intelligence: Theory, Algorithms, and Applications
- Benchmarking Instance-Centric Counterfactual Algorithms for XAI: From White Box to Black Box
- From Explanation to Recommendation: Ethical Standards for Algorithmic Recourse
- Decisions, Counterfactual Explanations and Strategic Behavior
- Interpretable Machine Learning: Moving From Mythos to Diagnostics
- Synthesizing explainable counterfactual policies for algorithmic recourse with program synthesis
- Improvement-Focused Causal Recourse (ICR)
- A Causal Perspective on Meaningful and Robust Algorithmic Recourse
- Setting the Right Expectations: Algorithmic Recourse Over Time
- Counterfactual Explanations for Arbitrary Regression Models
- The Cadaver in the Machine: The Social Practices of Measurement and Validation in Motion Capture Technology
- Optimal Counterfactual Explanations in Tree Ensembles
- Algorithmic Recourse in the Wild: Understanding the Impact of Data and Model Shifts
- Consistent Counterfactuals for Deep Models
- On the Connection between Game-Theoretic Feature Attributions and Counterfactual Explanations
- RoCourseNet: Distributionally Robust Training of a Prediction Aware Recourse Model
- The Importance of Time in Causal Algorithmic Recourse
- Understanding Prediction Discrepancies in Machine Learning Classifiers
- The Intriguing Relation Between Counterfactual Explanations and Adversarial Examples
- Exploring Counterfactual Explanations Through the Lens of Adversarial Examples: A Theoretical and Empirical Analysis
- Towards the Unification and Robustness of Perturbation and Gradient Based Explanations
- Ordered Counterfactual Explanation by Mixed-Integer Linear Optimization
- Counterfactual Explanations Can Be Manipulated
- Counterfactual Explanations for Machine Learning: Challenges Revisited
- Leveraging Sparse Linear Layers for Debuggable Deep Networks
- ABROCA Distributions For Algorithmic Bias Assessment: Considerations Around Interpretation
- From Model Performance to Claim: How a Change of Focus in Machine Learning Replicability Can Help Bridge the Responsibility Gap
- Navigating Explanatory Multiverse Through Counterfactual Path Geometry
- Towards Feasible Counterfactual Explanations: A Taxonomy Guided Template-based NLG Method
- Faithful and Plausible Explanations of Medical Code Predictions
- Perfect Counterfactuals in Imperfect Worlds: Modelling Noisy Implementation of Actions in Sequential Algorithmic Recourse
- Modeling Users and Online Communities for Abuse Detection: A Position on Ethics and Explainability
- A Series of Unfortunate Counterfactual Events: the Role of Time in Counterfactual Explanations