Empirical AUC for evaluating probabilistic forecasts
arXiv:1508.05503 · doi:10.1214/16-EJS1109
Abstract
Scoring functions are used to evaluate and compare partially probabilistic forecasts. We investigate the use of rank-sum functions such as empirical Area Under the Curve (AUC), a widely-used measure of classification performance, as a scoring function for the prediction of probabilities of a set of binary outcomes. It is shown that the AUC is not generally a proper scoring function, that is, under certain circumstances it is possible to improve on the expected AUC by modifying the quoted probabilities from their true values. However with some restrictions, or with certain modifications, it can be made proper.
15 pages
References in corpus (1)
Cited by in corpus (3)
- Performance evaluation and hyperparameter tuning of statistical and machine-learning models using spatial data
- CAGI, the Critical Assessment of Genome Interpretation, establishes progress and prospects for computational genetic variant interpretation methods
- Empirical AUC for evaluating probabilistic forecasts