Beyond temperature scaling: Obtaining well-calibrated multiclass probabilities with Dirichlet calibration
arXiv:1910.12656
Abstract
Class probabilities predicted by most multiclass classifiers are uncalibrated, often tending towards over-confidence. With neural networks, calibration can be improved by temperature scaling, a method to learn a single corrective multiplicative factor for inputs to the last softmax layer. On non-neural models the existing methods apply binary calibration in a pairwise or one-vs-rest fashion. We propose a natively multiclass calibration method applicable to classifiers from any model class, derived from Dirichlet distributions and generalising the beta calibration method from binary classification. It is easily implemented with neural nets since it is equivalent to log-transforming the uncalibrated probabilities, followed by one linear layer and softmax. Experiments demonstrate improved probabilistic predictions according to multiple measures (confidence-ECE, classwise-ECE, log-loss, Brier score) across a wide range of datasets and classifiers. Parameters of the learned Dirichlet calibration map provide insights to the biases in the uncalibrated model.
Accepted for presentation at NeurIPS 2019
Cited by in corpus (9)
- A Comparison of Uncertainty Estimation Approaches in Deep Learning Components for Autonomous Vehicle Applications
- Temporal Probability Calibration
- Approximating Instance-Dependent Noise via Instance-Confidence Embedding
- Quantile Regularization: Towards Implicit Calibration of Regression Models
- Calibrating Deep Neural Network Classifiers on Out-of-Distribution Datasets
- Spatially Varying Label Smoothing: Capturing Uncertainty from Expert Annotations
- Increasing Trustworthiness of Deep Neural Networks via Accuracy Monitoring
- Probabilistic Object Classification using CNN ML-MAP layers
- Learning to Cascade: Confidence Calibration for Improving the Accuracy and Computational Cost of Cascade Inference Systems