Consistency of random forests
arXiv:1405.2881 · doi:10.1214/15-AOS1321
Abstract
Random forests are a learning algorithm proposed by Breiman [Mach. Learn. 45 (2001) 5--32] that combines several randomized decision trees and aggregates their predictions by averaging. Despite its wide usage and outstanding practical performance, little is known about the mathematical properties of the procedure. This disparity between theory and practice originates in the difficulty to simultaneously analyze both the randomization process and the highly data-dependent tree structure. In the present paper, we take a step forward in forest exploration by proving a consistency result for Breiman's [Mach. Learn. 45 (2001) 5--32] original algorithm in the context of additive regression models. Our analysis also sheds an interesting light on how random forests can nicely adapt to sparsity. 1. Introduction. Random forests are an ensemble learning method for classification and regression that constructs a number of randomized decision trees during the training phase and predicts by averaging the results. Since its publication in the seminal paper of Breiman (2001), the procedure has become a major data analysis tool, that performs well in practice in comparison with many standard methods. What has greatly contributed to the popularity of forests is the fact that they can be applied to a wide range of prediction problems and have few parameters to tune. Aside from being simple to use, the method is generally recognized for its accuracy and its ability to deal with small sample sizes, high-dimensional feature spaces and complex data structures. The random forest methodology has been successfully involved in many practical problems, including air quality prediction (winning code of the EMC data science global hackathon in 2012, see http://www.kaggle.com/c/dsg-hackathon), chemoinformatics [Svetnik et al. (2003)], ecology [Prasad, Iverson and Liaw (2006), Cutler et al. (2007)], 3D
References in corpus (3)
Cited by in corpus (77)
- Meta-learners for Estimating Heterogeneous Treatment Effects using Machine Learning
- Unbiased split variable selection for random survival forests using maximally selected rank statistics
- Adaptive Concentration of Regression Trees, with Application to Random Forests
- High-dimensional regression adjustments in randomized experiments
- A Debiased MDI Feature Importance Measure for Random Forests
- Interpretable Random Forests via Rule Extraction
- Flexible domain prediction using mixed effects random forests
- Combining Static and Dynamic Features for Multivariate Sequence Classification
- Optimized ensemble deep learning framework for scalable forecasting of dynamics containing extreme events
- Selective inference for effect modification via the lasso
- Trees, forests, and impurity-based variable importance
- A novel distribution-free hybrid regression model for manufacturing process efficiency improvement
- Making Sense of Random Forest Probabilities: a Kernel Perspective
- Randomization as Regularization: A Degrees of Freedom Explanation for Random Forest Success
- Fast Linear Model Trees by PILOT
- Estimation and Inference with Trees and Forests in High Dimensions
- Scalable and Efficient Hypothesis Testing with Random Forests
- Getting Better from Worse: Augmented Bagging and a Cautionary Tale of Variable Importance
- SHAFF: Fast and consistent SHApley eFfect estimates via random Forests
- Unbiased Measurement of Feature Importance in Tree-Based Methods
- Provable Boolean Interaction Recovery from Tree Ensemble obtained via Random Forests
- Consistency of survival tree and forest models: splitting bias and correction
- The Implicit Regularization of Ordinary Least Squares Ensembles
- CyPhERS: A Cyber-Physical Event Reasoning System providing real-time situational awareness for attack and fault response
- Imputation procedures in surveys using nonparametric and machine learning methods: an empirical comparison
- Asymptotic Distributions and Rates of Convergence for Random Forests via Generalized U-statistics
- One Class Splitting Criteria for Random Forests
- Stochastic tree ensembles for regularized nonlinear regression
- Regression-Enhanced Random Forests
- Sharp Analysis of a Simple Model for Random Forests
- A Unified Framework for Random Forest Prediction Error Estimation
- Transformation Forests
- Adaptive Estimation of Multivariate Piecewise Polynomials and Bounded Variation Functions by Optimal Decision Trees
- Asymptotic Properties of High-Dimensional Random Forests
- Impact of subsampling and pruning on random forests
- On the asymptotics of random forests
- Extrapolated cross-validation for randomized ensembles
- -statistics and Variance Estimation
- Interpretable Classification of Bacterial Raman Spectra with Knockoff Wavelets
- Random Forests for dependent data
- A Practically Competitive and Provably Consistent Algorithm for Uplift Modeling
- Model-assisted estimation through random forests in finite population sampling
- GrCAN: Gradient Boost Convolutional Autoencoder with Neural Decision Forest
- Achieving Reliable Causal Inference with Data-Mined Variables: A Random Forest Approach to the Measurement Error Problem
- Maximizing Agreements for Ranking, Clustering and Hierarchical Clustering via MAX-CUT
- Subsampling Winner Algorithm for Feature Selection in Large Regression Data
- Reliable ABC model choice via random forests
- Semi-Supervised Off Policy Reinforcement Learning
- A cautionary tale on fitting decision trees to data from additive models: generalization lower bounds
- Random forests and kernel methods
- On the use of Harrell's C for clinical risk prediction via random survival forests
- A Numerical Transform of Random Forest Regressors corrects Systematically-Biased Predictions
- Causal survival embeddings: non-parametric counterfactual inference under censoring
- Nonparametric Variable Screening with Optimal Decision Stumps
- On the Consistency of a Random Forest Algorithm in the Presence of Missing Entries
- WildWood: a new Random Forest algorithm
- Classification Trees for Imbalanced and Sparse Data: Surface-to-Volume Regularization
- Building Trees for Probabilistic Prediction via Scoring Rules
- Adaptive Conformal Prediction by Reweighting Nonconformity Score
- Random Planted Forest: a directly interpretable tree ensemble
- Ensemble Learning with Statistical and Structural Models
- Best-scored Random Forest Classification
- Estimating the Algorithmic Variance of Randomized Ensembles via the Bootstrap
- Comments on Leo Breiman's paper 'Statistical Modeling: The Two Cultures' (Statistical Science, 2001, 16(3), 199-231)
- Modelling hetegeneous treatment effects by quantitle local polynomial decision tree and forest
- Robust Similarity and Distance Learning via Decision Forests
- Asymptotic Unbiasedness of the Permutation Importance Measure in Random Forest Models
- MMD-based Variable Importance for Distributional Random Forest
- Censored Quantile Regression Forests
- Two-stage Best-scored Random Forest for Large-scale Regression
- Wasserstein Random Forests and Applications in Heterogeneous Treatment Effects
- Sparse nearest neighbor Cholesky matrices in spatial statistics
- Computing Affine Combinations, Distances, and Correlations for Recursive Partition Functions
- Consistent Sufficient Explanations and Minimal Local Rules for explaining regression and classification models
- Bridging Breiman's Brook: From Algorithmic Modeling to Statistical Learning
- Modelling and Analysing Behaviours and Emotions via Complex User Interactions
- Prospective Learning in Retrospect