Forecast Evaluation of Quantiles, Prediction Intervals, and other Set-Valued Functionals
arXiv:1910.07912 · doi:10.1214/21-EJS1808
Abstract
We introduce a theoretical framework of elicitability and identifiability of set-valued functionals, such as quantiles, prediction intervals, and systemic risk measures. A functional is elicitable if it is the unique minimiser of an expected scoring function, and identifiable if it is the unique zero of an expected identification function; both notions are essential for forecast ranking and validation, and - and -estimation. Our framework distinguishes between exhaustive forecasts, being set-valued and aiming at correctly specifying the entire functional, and selective forecasts, content with solely specifying a single point in the correct functional. We establish a mutual exclusivity result: A set-valued functional can be either selectively elicitable or exhaustively elicitable or not elicitable at all. Notably, since quantiles are well known to be selectively elicitable, they fail to be exhaustively elicitable. We further show that the class of prediction intervals and Vorob'ev quantiles turn out to be exhaustively elicitable and selectively identifiable. In particular, we provide a mixture representation of elementary exhaustive scores, leading the way to Murphy diagrams. We give possibility and impossibility results for the shortest prediction interval and prediction intervals specified by an endpoint or a midpoint. We end with a comprehensive literature review on common practice in forecast evaluation of set-valued functionals.
46 pages, 2 figures. arXiv admin note: text overlap with arXiv:1907.01306
References in corpus (8)
- Evaluating epidemic forecasts in an interval format
- Scoring Interval Forecasts: Equal-Tailed, Shortest, and Modal Interval
- Why scoring functions cannot assess tail properties
- Monotone Least Squares and Isotonic Quantiles
- Supplement to "Erratum: Higher Order Elicitability and Osband's Principle"
- Elicitability and Identifiability of Systemic Risk Measures
- From Halfspace M-depth to Multiple-output Expectile Regression
- Calibration Scoring Rules for Practical Prediction Training
Cited by in corpus (10)
- A review of predictive uncertainty estimation with machine learning
- Sensitivity Measures Based on Scoring Functions
- Evaluating Range Value at Risk Forecasts
- Measurability of functionals and of ideal point forecasts
- Osband's Principle for Identification Functions
- Is the mode elicitable relative to unimodal distributions?
- Combinations of distributional regression algorithms with application in uncertainty estimation of corrected satellite precipitation products
- Ensemble learning for uncertainty estimation with application to the correction of satellite precipitation products
- Elicitability and identifiability of tail risk measures
- Variable transformations in consistent loss functions