Type I Error Rates are Not Usually Inflated
arXiv:2312.06265 · doi:10.36850/4d35-44bd
Abstract
The inflation of Type I error rates is thought to be one of the causes of the replication crisis. Questionable research practices such as p-hacking are thought to inflate Type I error rates above their nominal level, leading to unexpectedly high levels of false positives in the literature and, consequently, unexpectedly low replication rates. In this article, I offer an alternative view. I argue that questionable and other research practices do not usually inflate relevant Type I error rates. I begin by introducing the concept of Type I error rates and distinguishing between statistical errors and theoretical errors. I then illustrate my argument with respect to model misspecification, multiple testing, selective inference, forking paths, exploratory analyses, p-hacking, optional stopping, double dipping, and HARKing. In each case, I demonstrate that relevant Type I error rates are not usually inflated above their nominal level, and in the rare cases that they are, the inflation is easily identified and resolved. I conclude that the replication crisis may be explained, at least in part, by researchers' misinterpretation of statistical errors and their underestimation of theoretical errors.
References in corpus (6)
- When to adjust alpha during multiple testing: A consideration of disjunction, conjunction, and individual testing
- On the Birnbaum Argument for the Strong Likelihood Principle
- Where do statistical models come from? Revisiting the problem of specification
- Inconistent multiple testing corrections: The fallacy of using family-based error rates to make inferences about individual hypotheses
- Does preregistration improve the credibility of research findings?
- Connecting Simple and Precise P-values to Complex and Ambiguous Realities