Probabilistic performance estimators for computational chemistry methods: the empirical cumulative distribution function of absolute errors
arXiv:1801.03305 · doi:10.1063/1.5016248
Abstract
Benchmarking studies in computational chemistry use reference datasets to assess the accuracy of a method through error statistics. The commonly used error statistics, such as the mean signed and mean unsigned errors, do not inform end-users on the expected amplitude of prediction errors attached to these methods. We show that, the distributions of model errors being neither normal nor zero-centered, these error statistics cannot be used to infer prediction error probabilities. To overcome this limitation, we advocate for the use of more informative statistics, based on the empirical cumulative distribution function of unsigned errors, namely (1) the probability for a new calculation to have an absolute error below a chosen threshold, and (2) the maximal amplitude of errors one can expect with a chosen high confidence level. Those statistics are also shown to be well suited for benchmarking and ranking studies. Moreover, the standard error on all benchmarking statistics depends on the size of the reference dataset. Systematic publication of these standard errors would be very helpful to assess the statistical reliability of benchmarking conclusions.
Supplementary material: https://github.com/ppernot/ECDFT
References in corpus (3)
Cited by in corpus (14)
- Guest Editorial: Special Topic on Data-enabled Theoretical Chemistry
- Universal QM/MM Approaches for General Nanoscale Applications
- Self-Parametrizing System-Focused Atomistic Models
- The long road to calibrated prediction uncertainty in computational chemistry
- Statistical Analysis of Semiclassical Dispersion Corrections
- Molecule-specific Uncertainty Quantification in Quantum Chemical Studies
- Impact of non-normal error distributions on the benchmarking and ranking of Quantum Machine Learning models
- Probabilistic performance estimators for computational chemistry methods: Systematic Improvement Probability and Ranking Probability Matrix. I. Theory
- Prediction uncertainty validation for computational chemists
- Critical Benchmarking of the G4(MP2) Model, the Correlation Consistent Composite Approach and Popular Density Functional Approximations on a Probabilistically Pruned Benchmark Dataset of Formation Enthalpies
- Using the Gini coefficient to characterize the shape of computational chemistry error distributions
- Probabilistic performance estimators for computational chemistry methods: Systematic Improvement Probability and Ranking Probability Matrix. II. Applications
- Towards Ultra Low Cobalt Cathodes: A High Fidelity Computational Phase Search of Layered Li-Ni-Mn-Co Oxides
- Model selection in atomistic simulation