Model-agnostic Feature Importance and Effects with Dependent Features -- A Conditional Subgroup Approach
arXiv:2006.04628 · doi:10.1007/s10618-022-00901-9
Abstract
The interpretation of feature importance in machine learning models is challenging when features are dependent. Permutation feature importance (PFI) ignores such dependencies, which can cause misleading interpretations due to extrapolation. A possible remedy is more advanced conditional PFI approaches that enable the assessment of feature importance conditional on all other features. Due to this shift in perspective and in order to enable correct interpretations, it is therefore important that the conditioning is transparent and humanly comprehensible. In this paper, we propose a new sampling mechanism for the conditional distribution based on permutations in conditional subgroups. As these subgroups are constructed using decision trees (transformation trees), the conditioning becomes inherently interpretable. This not only provides a simple and effective estimator of conditional PFI, but also local PFI estimates within the subgroups. In addition, we apply the conditional subgroups approach to partial dependence plots (PDP), a popular method for describing feature effects that can also suffer from extrapolation when features are dependent and interactions are present in the model. We show that PFI and PDP based on conditional subgroups often outperform methods such as conditional PFI based on knockoffs, or accumulated local effect plots. Furthermore, our approach allows for a more fine-grained interpretation of feature effects and importance within the conditional subgroups.
References in corpus (3)
Cited by in corpus (14)
- Interpretable Machine Learning -- A Brief History, State-of-the-Art and Challenges
- Adversarial attacks and defenses in explainable artificial intelligence: A survey
- Relating the Partial Dependence Plot and Permutation Feature Importance to the Data Generating Process
- Grouped Feature Importance and Combined Features Effect Plot
- Incremental Permutation Feature Importance (iPFI): Towards Online Explanations on Data Streams
- A Guide to Feature Importance Methods for Scientific Inference
- Scientific Inference With Interpretable Machine Learning: Analyzing Models to Learn About Real-World Phenomena
- Opening the random forest black box by the analysis of the mutual impact of features
- Transforming Feature Space to Interpret Machine Learning Models
- Beyond development: Challenges in deploying machine learning models for structural engineering applications
- Postprocessing of Ensemble Weather Forecasts Using Permutation-invariant Neural Networks
- Interpretable machine learning for time-to-event prediction in medicine and healthcare
- Conditional Feature Importance for Mixed Data
- Automated Speech Scoring System Under The Lens: Evaluating and interpreting the linguistic cues for language proficiency