Calibration in Machine Learning Uncertainty Quantification: beyond consistency to target adaptivity
arXiv:2309.06240 · doi:10.1063/5.0174943
Abstract
Reliable uncertainty quantification (UQ) in machine learning (ML) regression tasks is becoming the focus of many studies in materials and chemical science. It is now well understood that average calibration is insufficient, and most studies implement additional methods testing the conditional calibration with respect to uncertainty, i.e. consistency. Consistency is assessed mostly by so-called reliability diagrams. There exists however another way beyond average calibration, which is conditional calibration with respect to input features, i.e. adaptivity. In practice, adaptivity is the main concern of the final users of a ML-UQ method, seeking for the reliability of predictions and uncertainties for any point in features space. This article aims to show that consistency and adaptivity are complementary validation targets, and that a good consistency does not imply a good adaptivity. Adapted validation methods are proposed and illustrated on a representative example.
arXiv admin note: text overlap with arXiv:2303.07170
References in corpus (10)
- On Calibration of Modern Neural Networks
- Uncertainty Toolbox: an Open-Source Library for Assessing, Visualizing, and Improving Uncertainty Quantification
- The long road to calibrated prediction uncertainty in computational chemistry
- Individual Calibration with Randomized Forecasting
- The Peril of Popular Deep Learning Uncertainty Estimation Methods
- A Farewell to the Bias-Variance Tradeoff? An Overview of the Theory of Overparameterized Machine Learning
- Improving Conditional Coverage via Orthogonal Quantile Regression
- Validation of uncertainty quantification metrics: a primer based on the consistency and adaptivity concepts
- Properties of the ENCE and other MAD-based calibration metrics
- Stratification of uncertainties recalibrated by isotonic regression and its impact on calibration error statistics