Counterfactual Shapley Additive Explanations
arXiv:2110.14270 · doi:10.1145/3531146.3533168
Abstract
Feature attributions are a common paradigm for model explanations due to their simplicity in assigning a single numeric score for each input feature to a model. In the actionable recourse setting, wherein the goal of the explanations is to improve outcomes for model consumers, it is often unclear how feature attributions should be correctly used. With this work, we aim to strengthen and clarify the link between actionable recourse and feature attributions. Concretely, we propose a variant of SHAP, Counterfactual SHAP (CF-SHAP), that incorporates counterfactual information to produce a background dataset for use within the marginal (a.k.a. interventional) Shapley value framework. We motivate the need within the actionable recourse setting for careful consideration of background datasets when using Shapley values for feature attributions with numerous synthetic examples. Moreover, we demonstrate the efficacy of CF-SHAP by proposing and justifying a quantitative score for feature attributions, counterfactual-ability, showing that as measured by this metric, CF-SHAP is superior to existing methods when evaluated on public datasets using tree ensembles.
Accepted at FAccT '22 (2022 ACM Conference on Fairness, Accountability, and Transparency)
References in corpus (10)
- The Hidden Assumptions Behind Counterfactual Explanations and Principal Reasons
- Feature relevance quantification in explainable AI: A causal problem
- Towards Unifying Feature Attribution and Counterfactual Explanations: Different Means to the Same End
- Generating Counterfactual and Contrastive Explanations using SHAP
- True to the Model or True to the Data?
- PermuteAttack: Counterfactual Explanation of Machine Learning Credit Scorecards
- Shapley Flow: A Graph-based Approach to Interpreting Model Predictions
- Beyond Individualized Recourse: Interpretable and Interactive Summaries of Actionable Recourses
- Rational Shapley Values
- Counterfactual Explanations for Arbitrary Regression Models