Uncertainty Quantification Metrics for Deep Regression
arXiv:2405.04278 · doi:10.1016/j.patrec.2024.09.011
Abstract
When deploying deep neural networks on robots or other physical systems, the learned model should reliably quantify predictive uncertainty. A reliable uncertainty allows downstream modules to reason about the safety of its actions. In this work, we address metrics for evaluating such an uncertainty. Specifically, we focus on regression tasks, and investigate Area Under Sparsification Error (AUSE), Calibration Error, Spearman's Rank Correlation, and Negative Log-Likelihood (NLL). Using synthetic regression datasets, we look into how those metrics behave under four typical types of uncertainty, their stability regarding the size of the test set, and reveal their strengths and weaknesses. Our results indicate that Calibration Error is the most stable and interpretable metric, but AUSE and NLL also have their respective use cases. We discourage the usage of Spearman's Rank Correlation for evaluating uncertainties and recommend replacing it with AUSE.
References in corpus (5)
- Single-model uncertainty quantification in neural network potentials does not consistently outperform model ensembles
- Estimating Uncertainty in Neural Networks for Cardiac MRI Segmentation: A Benchmark Study
- Wasserstein Distances for Stereo Disparity Estimation
- A Simple Approach to Improve Single-Model Deep Uncertainty via Distance-Awareness
- Learning Sample Difficulty from Pre-trained Models for Reliable Prediction