Statistical methods for linguistic research: Foundational Ideas - Part I
arXiv:1601.01126 · doi:10.1111/lnc3.12201
Abstract
We present the fundamental ideas underlying statistical hypothesis testing using the frequentist framework. We begin with a simple example that builds up the one-sample t-test from the beginning, explaining important concepts such as the sampling distribution of the sample mean, and the iid assumption. Then we examine the p-value in detail, and discuss several important misconceptions about what a p-value does and does not tell us. This leads to a discussion of Type I, II error and power, and Type S and M error. An important conclusion from this discussion is that one should aim to carry out appropriately powered studies. Next, we discuss two common issues we have encountered in psycholinguistics and linguistics: running experiments until significance is reached, and the "garden-of-forking-paths" problem discussed by Gelman and others, whereby the researcher attempts to find statistical significance by analyzing the data in different ways. The best way to use frequentist methods is to run appropriately powered studies, check model assumptions, clearly separate exploratory data analysis from confirmatory hypothesis testing, and always attempt to replicate results.
30 pages, 9 figures, 3 tables. Under review with Language and Linguistics Compass. Comments and suggestions for improvement welcome. (Replaced version corrects several typos)
References in corpus (6)
- Asymptotic Equivalence of Bayes Cross Validation and Widely Applicable Information Criterion in Singular Learning Theory
- Balancing Type I Error and Power in Linear Mixed Models
- Parsimonious Mixed Models
- A General Framework for the Parametrization of Hierarchical Models
- Scientific Utopia: II. Restructuring incentives and practices to promote truth over publishability
- Bayesian linear mixed models using Stan: A tutorial for psychologists, linguists, and cognitive scientists