PredDiff: Explanations and Interactions from Conditional Expectations
arXiv:2102.13519 · doi:10.1016/j.artint.2022.103774
Abstract
PredDiff is a model-agnostic, local attribution method that is firmly rooted in probability theory. Its simple intuition is to measure prediction changes while marginalizing features. In this work, we clarify properties of PredDiff and its close connection to Shapley values. We stress important differences between classification and regression, which require a specific treatment within both formalisms. We extend PredDiff by introducing a new, well-founded measure for interaction effects between arbitrary feature subsets. The study of interaction effects represents an inevitable step towards a comprehensive understanding of black-box models and is particularly important for science applications. Equipped with our novel interaction measure, PredDiff is a promising model-agnostic approach for obtaining reliable, numerically inexpensive and theoretically sound attributions.
35 pages, 20 Figures, accepted journal version, code available at https://github.com/AI4HealthUOL/preddiff-interactions
References in corpus (12)
- Methods for Interpreting and Understanding Deep Neural Networks
- On Calibration of Modern Neural Networks
- Unmasking Clever Hans Predictors and Assessing What Machines Really Learn
- Visualizing Deep Neural Network Decisions: Prediction Difference Analysis
- Feature relevance quantification in explainable AI: A causal problem
- Explaining by Removing: A Unified Framework for Model Explanation
- Toward Explainable AI for Regression Models
- Benchmarking Deep Learning Interpretability in Time Series Predictions
- Quantifying and Visualizing Attribute Interactions
- Fairwashing Explanations with Off-Manifold Detergent
- How does this interaction affect me? Interpretable attribution for feature interactions
- Explaining Time Series Predictions with Dynamic Masks
Cited by in corpus (4)
- From Attribution Maps to Human-Understandable Explanations through Concept Relevance Propagation
- An XAI framework for robust and transparent data-driven wind turbine power curve models
- Insights Into the Inner Workings of Transformer Models for Protein Function Prediction
- Towards Symbolic XAI -- Explanation Through Human Understandable Logical Relationships Between Features