The Theory Behind Overfitting, Cross Validation, Regularization, Bagging, and Boosting: Tutorial
arXiv:1905.12787
Abstract
In this tutorial paper, we first define mean squared error, variance, covariance, and bias of both random variables and classification/predictor models. Then, we formulate the true and generalization errors of the model for both training and validation/test instances where we make use of the Stein's Unbiased Risk Estimator (SURE). We define overfitting, underfitting, and generalization using the obtained true and generalization errors. We introduce cross validation and two well-known examples which are -fold and leave-one-out cross validations. We briefly introduce generalized cross validation and then move on to regularization where we use the SURE again. We work on both and norm regularizations. Then, we show that bootstrap aggregating (bagging) reduces the variance of estimation. Boosting, specifically AdaBoost, is introduced and it is explained as both an additive model and a maximum margin model, i.e., Support Vector Machine (SVM). The upper bound on the generalization error of boosting is also provided to show why boosting prevents from overfitting. As examples of regularization, the theory of ridge and lasso regressions, weight decay, noise injection to input/weights, and early stopping are explained. Random forest, dropout, histogram of oriented gradients, and single shot multi-box detector are explained as examples of bagging in machine learning and computer vision. Finally, boosting tree and SVM models are mentioned as examples of boosting.
23 pages, 9 figures. v2: typos are fixed
Cited by in corpus (12)
- Comparison between different methods of model selection in cosmology
- Isolation Mondrian Forest for Batch and Online Anomaly Detection
- KKT Conditions, First-Order and Second-Order Optimization, and Distributed Optimization: Tutorial and Survey
- Fisher and Kernel Fisher Discriminant Analysis: Tutorial
- Prediction and understanding of soft proton contamination in XMM-Newton: a machine learning approach
- Sampling Algorithms, from Survey Sampling to Monte Carlo Methods: Tutorial and Literature Review
- Johnson-Lindenstrauss Lemma, Linear and Nonlinear Random Projections, Random Fourier Features, and Random Kitchen Sinks: Tutorial and Survey
- Generative Adversarial Networks and Adversarial Autoencoders: Tutorial and Survey
- Backprojection for Training Feedforward Neural Networks in the Input and Feature Spaces
- Restricted Boltzmann Machine and Deep Belief Network: Tutorial and Survey
- Acceleration of Large Margin Metric Learning for Nearest Neighbor Classification Using Triplet Mining and Stratified Sampling
- Cascade Bagging for Accuracy Prediction with Few Training Samples