Time for a change: a tutorial for comparing multiple classifiers through Bayesian analysis
arXiv:1606.04316
Abstract
The machine learning community adopted the use of null hypothesis significance testing (NHST) in order to ensure the statistical validity of results. Many scientific fields however realized the shortcomings of frequentist reasoning and in the most radical cases even banned its use in publications. We should do the same: just as we have embraced the Bayesian paradigm in the development of new machine learning methods, so we should also use it in the analysis of our own results. We argue for abandonment of NHST by exposing its fallacies and, more importantly, offer better - more sound and useful - alternatives for it.
This paper has been published in the Journal of Machine Learning Research (JMLR) vol.18, 2017
References in corpus (1)
Cited by in corpus (7)
- Evaluating time series forecasting models: An empirical study on performance estimation methods
- Retweet communities reveal the main sources of hate speech
- Propositionalization and Embeddings: Two Sides of the Same Coin
- Predicting Acute Kidney Injury at Hospital Re-entry Using High-dimensional Electronic Health Record Data
- Deep Node Ranking for Neuro-symbolic Structural Node Embedding and Classification
- Meta-Learning for Unsupervised Outlier Detection with Optimal Transport
- Using Machine Learning and Natural Language Processing Techniques to Analyze and Support Moderation of Student Book Discussions