Empirical Frequentist Coverage of Deep Learning Uncertainty Quantification Procedures
arXiv:2010.03039 · doi:10.3390/e23121608
Abstract
Uncertainty quantification for complex deep learning models is increasingly important as these techniques see growing use in high-stakes, real-world settings. Currently, the quality of a model's uncertainty is evaluated using point-prediction metrics such as negative log-likelihood or the Brier score on heldout data. In this study, we provide the first large scale evaluation of the empirical frequentist coverage properties of well known uncertainty quantification techniques on a suite of regression and classification tasks. We find that, in general, some methods do achieve desirable coverage properties on in distribution samples, but that coverage is not maintained on out-of-distribution data. Our results demonstrate the failings of current uncertainty quantification techniques as dataset shift increases and establish coverage as an important metric in developing models for real-world applications.
13 pages, 13 figures
References in corpus (13)
- On Calibration of Modern Neural Networks
- Can You Trust Your Model's Uncertainty? Evaluating Predictive Uncertainty Under Dataset Shift
- Simple and Scalable Predictive Uncertainty Estimation using Deep Ensembles
- Probabilistic Backpropagation for Scalable Learning of Bayesian Neural Networks
- Uncertainty Estimation Using a Single Deep Deterministic Neural Network
- Multiplicative Normalizing Flows for Variational Bayesian Neural Networks
- Flipout: Efficient Pseudo-Independent Weight Perturbations on Mini-Batches
- A Simple Baseline for Bayesian Uncertainty in Deep Learning
- Simple and Principled Uncertainty Estimation with Deterministic Deep Learning via Distance Awareness
- Quality of Uncertainty Quantification for Bayesian Neural Network Inference
- Implicit Weight Uncertainty in Neural Networks
- How Good is the Bayes Posterior in Deep Neural Networks Really?
- Deep Bayesian Bandits Showdown: An Empirical Comparison of Bayesian Deep Networks for Thompson Sampling