A Guide to Feature Importance Methods for Scientific Inference
arXiv:2404.12862 · doi:10.1007/978-3-031-63797-1_22
Abstract
While machine learning (ML) models are increasingly used due to their high predictive power, their use in understanding the data-generating process (DGP) is limited. Understanding the DGP requires insights into feature-target associations, which many ML models cannot directly provide due to their opaque internal mechanisms. Feature importance (FI) methods provide useful insights into the DGP under certain conditions. Since the results of different FI methods have different interpretations, selecting the correct FI method for a concrete use case is crucial and still requires expert knowledge. This paper serves as a comprehensive guide to help understand the different interpretations of global FI methods. Through an extensive review of FI methods and providing new proofs regarding their interpretation, we facilitate a thorough understanding of these methods and formulate concrete recommendations for scientific inference. We conclude by discussing options for FI uncertainty estimation and point to directions for future research aiming at full statistical inference from black-box ML models.
References in corpus (8)
- To Explain or to Predict?
- The Hardness of Conditional Independence Testing and the Generalised Covariance Measure
- Relating the Partial Dependence Plot and Permutation Feature Importance to the Data Generating Process
- Model-agnostic Feature Importance and Effects with Dependent Features -- A Conditional Subgroup Approach
- OpenXAI: Towards a Transparent Evaluation of Model Explanations
- A general framework for inference on algorithm-agnostic variable importance
- Grouped Feature Importance and Combined Features Effect Plot
- Conditional Feature Importance for Mixed Data