Verified Uncertainty Calibration
arXiv:1909.10155
Abstract
Applications such as weather forecasting and personalized medicine demand models that output calibrated probability estimates---those representative of the true likelihood of a prediction. Most models are not calibrated out of the box but are recalibrated by post-processing model outputs. We find in this work that popular recalibration methods like Platt scaling and temperature scaling are (i) less calibrated than reported, and (ii) current techniques cannot estimate how miscalibrated they are. An alternative method, histogram binning, has measurable calibration error but is sample inefficient---it requires samples, compared to for scaling methods, where is the number of distinct probabilities the model can output. To get the best of both worlds, we introduce the scaling-binning calibrator, which first fits a parametric function to reduce variance and then bins the function values to actually ensure calibration. This requires only samples. Next, we show that we can estimate a model's calibration error more accurately using an estimator from the meteorological community---or equivalently measure its calibration error with fewer samples ( instead of ). We validate our approach with multiclass calibration experiments on CIFAR-10 and ImageNet, where we obtain a 35% lower calibration error than histogram binning and, unlike scaling methods, guarantees on true calibration. In these experiments, we also estimate the calibration error and ECE more accurately than the commonly used plugin estimators. We implement all these methods in a Python library: https://pypi.org/project/uncertainty-calibration
Accepted as a spotlight to NeurIPS 2019, updated to include experiments for ECE
References in corpus (7)
- On Calibration of Modern Neural Networks
- Deep Anomaly Detection with Outlier Exposure
- Beyond temperature scaling: Obtaining well-calibrated multiclass probabilities with Dirichlet calibration
- Measuring Calibration in Deep Learning
- Evaluating model calibration in classification
- Calibration tests in multi-class classification: A unifying framework
- Calibrated Model-Based Deep Reinforcement Learning
Cited by in corpus (19)
- A Review of Uncertainty Quantification in Deep Learning: Techniques, Applications and Challenges
- A Review and Comparative Study on Probabilistic Object Detection in Autonomous Driving
- Evaluating probabilistic classifiers: Reliability diagrams and score decompositions revisited
- MoPro: Webly Supervised Learning with Momentum Prototypes
- Calibrating AI Models for Few-Shot Demodulation via Conformal Prediction
- On the Validity of Bayesian Neural Networks for Uncertainty Estimation
- Cal-SFDA: Source-Free Domain-adaptive Semantic Segmentation with Differentiable Expected Calibration Error
- Don't Just Blame Over-parametrization for Over-confidence: Theoretical Analysis of Calibration in Binary Classification
- CRUDE: Calibrating Regression Uncertainty Distributions Empirically
- Assurance Monitoring of Learning Enabled Cyber-Physical Systems Using Inductive Conformal Prediction based on Distance Learning
- On the Role of Dataset Quality and Heterogeneity in Model Confidence
- Exploring Covariate and Concept Shift for Detection and Calibration of Out-of-Distribution Data
- Post-hoc Models for Performance Estimation of Machine Learning Inference
- Confidence Calibration with Bounded Error Using Transformations
- Improving Uncertainty-Error Correspondence in Deep Bayesian Medical Image Segmentation
- Understanding the Under-Coverage Bias in Uncertainty Estimation
- Making Heads and Tails of Models with Marginal Calibration for Sparse Tagsets
- Uncertainty Propagation in Node Classification
- Towards Robust Active Feature Acquisition