Better Uncertainty Calibration via Proper Scores for Classification and Beyond
arXiv:2203.07835
Abstract
With model trustworthiness being crucial for sensitive real-world applications, practitioners are putting more and more focus on improving the uncertainty calibration of deep neural networks. Calibration errors are designed to quantify the reliability of probabilistic predictions but their estimators are usually biased and inconsistent. In this work, we introduce the framework of proper calibration errors, which relates every calibration error to a proper score and provides a respective upper bound with optimal estimation properties. This relationship can be used to reliably quantify the model calibration improvement. We theoretically and empirically demonstrate the shortcomings of commonly used estimators compared to our approach. Due to the wide applicability of proper scores, this gives a natural extension of recalibration beyond classification.
Published at NeurIPS 2022. Corrected conference version Theorem 3.1 and Proposition 3.2 since CWCE=0 does not imply TCE=0
Cited by in corpus (4)
- Understanding metric-related pitfalls in image analysis validation
- Comparing hundreds of machine learning classifiers and discrete choice models in predicting travel behavior: an empirical benchmark
- SAUC: Sparsity-Aware Uncertainty Calibration for Spatiotemporal Prediction with Graph Neural Networks
- Improving Uncertainty-Error Correspondence in Deep Bayesian Medical Image Segmentation