Obtaining Calibrated Probabilities from Boosting
arXiv:1207.1403
Abstract
Boosted decision trees typically yield good accuracy, precision, and ROC area. However, because the outputs from boosting are not well calibrated posterior probabilities, boosting yields poor squared error and cross-entropy. We empirically demonstrate why AdaBoost predicts distorted probabilities and examine three calibration methods for correcting this distortion: Platt Scaling, Isotonic Regression, and Logistic Correction. We also experiment with boosting using log-loss instead of the usual exponential loss. Experiments show that Logistic Correction and boosting with log-loss work well when boosting weak models such as decision stumps, but yield poor performance when boosting more complex models such as full decision trees. Platt Scaling and Isotonic Regression, however, significantly improve the probabilities predicted by
Appears in Proceedings of the Twenty-First Conference on Uncertainty in Artificial Intelligence (UAI2005)
Cited by in corpus (13)
- Multi-View Self-Attention for Interpretable Drug-Target Interaction Prediction
- CCMI : Classifier based Conditional Mutual Information Estimation
- Evidence Networks: simple losses for fast, amortized, neural Bayesian model comparison
- ReMix: Calibrated Resampling for Class Imbalance in Deep learning
- Leveraging Uncertainty in Deep Learning for Selective Classification
- Loss Functions for Discrete Contextual Pricing with Observational Data
- Misclassification cost-sensitive ensemble learning: A unifying framework
- Combining Forecasts Using Ensemble Learning
- Refinement revisited with connections to Bayes error, conditional entropy and calibrated classifiers
- Hands-Free Segmentation of Medical Volumes via Binary Inputs
- Calibration for Stratified Classification Models
- Better Boosting with Bandits for Online Learning
- Calibrated Boosting-Forest