Correlation and variable importance in random forests
arXiv:1310.5726 · doi:10.1007/s11222-016-9646-1
Abstract
This paper is about variable selection with the random forests algorithm in presence of correlated predictors. In high-dimensional regression or classification frameworks, variable selection is a difficult task, that becomes even more challenging in the presence of highly correlated predictors. Firstly we provide a theoretical study of the permutation importance measure for an additive regression model. This allows us to describe how the correlation between predictors impacts the permutation importance. Our results motivate the use of the Recursive Feature Elimination (RFE) algorithm for variable selection in this context. This algorithm recursively eliminates the variables using permutation importance measure as a ranking criterion. Next various simulation experiments illustrate the efficiency of the RFE algorithm for selecting a small number of variables together with a good prediction error. Finally, this selection algorithm is tested on the Landsat Satellite data from the UCI Machine Learning Repository.
References in corpus (3)
Cited by in corpus (33)
- All Models are Wrong, but Many are Useful: Learning a Variable's Importance by Studying an Entire Class of Prediction Models Simultaneously
- Machine Learning Predicts Laboratory Earthquakes
- Grouped variable importance with random forests and application to multiple functional data analysis
- Visualizing the Feature Importance for Black Box Models
- Identifying Solar Flare Precursors Using Time Series of SDO/HMI Images and SHARP Parameters
- Model-agnostic Feature Importance and Effects with Dependent Features -- A Conditional Subgroup Approach
- A comparative study of feature selection methods for stress hotspot classification in materials
- Incremental Permutation Feature Importance (iPFI): Towards Online Explanations on Data Streams
- Boosting Random Forests to Reduce Bias; One-Step Boosted Forest and its Variance Estimate
- Trees, forests, and impurity-based variable importance
- Feature Inference Attack on Shapley Values
- LEAF: Navigating Concept Drift in Cellular Networks
- Unbiased Measurement of Feature Importance in Tree-Based Methods
- FedScore: A privacy-preserving framework for federated scoring system development
- Two-Stage Human Verification using HandCAPTCHA and Anti-Spoofed Finger Biometrics with Feature Selection
- Concept Tree: High-Level Representation of Variables for More Interpretable Surrogate Decision Trees
- Approximate False Positive Rate Control in Selection Frequency for Random Forest
- Estimating the Physical State of a Laboratory Slow Slipping Fault from Seismic Signals
- WATCH: A Workflow to Assess Treatment Effect Heterogeneity in Drug Development for Clinical Trial Sponsors
- A Subspace-based Approach for Dimensionality Reduction and Important Variable Selection
- Importance measures derived from random forests: characterisation and extension
- Layer-based Composite Reputation Bootstrapping
- Predicting laboratory earthquakes with machine learning
- Flexible Predictive Distributions from Varying-Thresholds Modelling
- Interpretable random forest models through forward variable selection
- Sequential Feature Classification in the Context of Redundancies
- SalienTrack: providing salient information for semi-automated self-tracking feedback with model explanations
- Asymptotic Normality for Multivariate Random Forest Estimators
- The Threat to the Validity of Predictive Mutation Testing: The Impact of Uncovered Mutants
- Asymptotic Unbiasedness of the Permutation Importance Measure in Random Forest Models
- Machine learning pipeline for battery state of health estimation
- MMD-based Variable Importance for Distributional Random Forest
- Morpheus: Lightweight RTT Prediction for Performance-Aware Load Balancing