Permutation p-values should never be zero: calculating exact p-values when permutations are randomly drawn
arXiv:1603.05766 · doi:10.2202/1544-6115.1585
Abstract
Permutation tests are amongst the most commonly used statistical tools in modern genomic research, a process by which p-values are attached to a test statistic by randomly permuting the sample or gene labels. Yet permutation p-values published in the genomic literature are often computed incorrectly, understated by about 1/m, where m is the number of permutations. The same is often true in the more general situation when Monte Carlo simulation is used to assign p-values. Although the p-value understatement is usually small in absolute terms, the implications can be serious in a multiple testing context. The understatement arises from the intuitive but mistaken idea of using permutation to estimate the tail probability of the test statistic. We argue instead that permutation should be viewed as generating an exact discrete null distribution. The relevant literature, some of which is likely to have been relatively inaccessible to the genomic community, is reviewed and summarized. A computation strategy is developed for exact p-values when permutations are randomly drawn. The strategy is valid for any number of permutations and samples. Some simple recommendations are made for the implementation of permutation tests in practice.
12 pages, 2 figures
Cited by in corpus (23)
- Is K-fold cross validation the best model selection method for Machine Learning?
- Fast Approximation of Small p-values in Permutation Tests by Partitioning the Permutations
- Another look at the Lady Tasting Tea and differences between permutation tests and randomization tests
- New methods for multiple testing in permutation inference for the general linear model
- Scalable and Efficient Hypothesis Testing with Random Forests
- Two-sample goodness-of-fit tests on the flat torus based on Wasserstein distance and their relevance to structural biology
- Robust Functional ANOVA with Application to Additive Manufacturing
- Toward a test of Gaussianity of a gravitational wave background
- Statistical Agnostic Regression: a machine learning method to validate regression models
- Automatic Detection of Myocontrol Failures Based upon Situational Context Information
- An in silico drug repurposing pipeline to identify drugs with the potential to inhibit SARS-CoV-2 replication
- Adaptive Monte Carlo Multiple Testing via Multi-Armed Bandits
- Functional Inference on Rotational Curves and Identification of Human Gait at the Knee Joint
- Precise radial velocities of giant stars -- XVII. Distinguishing planets from intrinsically induced radial velocity signals in evolved stars
- Prodromal Diagnosis of Lewy Body Diseases Based on the Assessment of Graphomotor and Handwriting Difficulties
- Impeding Turbulence Decay in Self-gravitating Cloud Cores
- Extrema-weighted feature extraction for functional data
- Conditional variational autoencoders for cosmological model discrimination and anomaly detection in cosmic microwave background power spectra
- A Distance-Based Test of Association Between Paired Heterogeneous Genomic Data
- FDR Control for Online Anomaly Detection
- Stable Feature Selection with Applications to MALDI Imaging Mass Spectrometry Data
- Fourier-domain transfer entropy spectrum
- The Dynamics of Human and AI-Generated Language: How Semantics Fluctuates across Different Timescales