The Evolution of Boosting Algorithms - From Machine Learning to Statistical Modelling
arXiv:1403.1452 · doi:10.3414/ME13-01-0122
Abstract
The concept of boosting emerged from the field of machine learning. The basic idea is to boost the accuracy of a weak classifying tool by combining various instances into a more accurate prediction. This general concept was later adapted to the field of statistical modelling. This review article attempts to highlight this evolution of boosting algorithms from machine learning to statistical modelling. We describe the AdaBoost algorithm for classification as well as the two most prominent statistical boosting approaches, gradient boosting and likelihood-based boosting. Although both appraoches are typically treated separately in the literature, they share the same methodological roots and follow the same fundamental concepts. Compared to the initial machine learning algorithms, which must be seen as black-box prediction schemes, statistical boosting result in statistical models which offer a straight-forward interpretation. We highlight the methodological background and present the most common software implementations. Worked out examples and corresponding R code can be found in the Appendix.
References in corpus (5)
- Boosting Algorithms: Regularization, Prediction and Model Fitting
- Boosting the concordance index for survival data - a unified framework to derive and evaluate biomarker combinations
- Extending Statistical Boosting - An Overview of Recent Methodological Developments
- Comment: Boosting Algorithms: Regularization, Prediction and Model Fitting
- Rejoinder: Boosting Algorithms: Regularization, Prediction and Model Fitting
Cited by in corpus (17)
- A review of predictive uncertainty estimation with machine learning
- Super ensemble learning for daily streamflow forecasting: Large-scale demonstration and comparison with multiple machine learning algorithms
- Boosting algorithms in energy research: A systematic review
- Stability selection for component-wise gradient boosting in multiple dimensions
- Separation of pulsar signals from noise with supervised machine learning algorithms
- Extending Statistical Boosting - An Overview of Recent Methodological Developments
- Machine learning classification of CHIME fast radio bursts -- I. Supervised methods
- Merging satellite and gauge-measured precipitation using LightGBM with an emphasis on extreme quantiles
- Comparison of tree-based ensemble algorithms for merging satellite and earth-observed precipitation data at the daily time scale
- A Comparative Study of Methods for Estimating Conditional Shapley Values and When to Use Them
- Comparison of machine learning algorithms for merging gridded satellite and earth-observed precipitation data
- Contrast Trees and Distribution Boosting
- Probabilistic water demand forecasting using quantile regression algorithms
- Ensemble learning for blending gridded satellite and gauge-measured precipitation data
- Uncertainty estimation of machine learning spatial precipitation predictions from satellite data
- Combinations of distributional regression algorithms with application in uncertainty estimation of corrected satellite precipitation products
- Ensemble learning for uncertainty estimation with application to the correction of satellite precipitation products