Waste Not, Want Not: Why Rarefying Microbiome Data is Inadmissible
arXiv:1310.0424 · doi:10.1371/journal.pcbi.1003531
Abstract
The interpretation of count data originating from the current generation of DNA sequencing platforms requires special attention. In particular, the per-sample library sizes often vary by orders of magnitude from the same sequencing run, and the counts are overdispersed relative to a simple Poisson model These challenges can be addressed using an appropriate mixture model that simultaneously accounts for library size differences and biological variability. This approach is already well-characterized and implemented for RNA-Seq data in R packages such as edgeR and DESeq. We use statistical theory, extensive simulations, and empirical data to show that variance stabilizing normalization using a mixture model like the negative binomial is appropriate for microbiome count data. In simulations detecting differential abundance, normalization procedures based on a Gamma-Poisson mixture model provided systematic improvement in performance over crude proportions or rarefied counts -- both of which led to a high rate of false positives. In simulations evaluating clustering accuracy, we found that the rarefying procedure discarded samples that were nevertheless accurately clustered by alternative methods, and that the choice of minimum library size threshold was critical in some settings, but with an optimum that is unknown in practice. Techniques that use variance stabilizing transformations by modeling microbiome count data with a mixture distribution, such as those implemented in edgeR and DESeq, substantially improved upon techniques that attempt to normalize by rarefying or crude proportions. Based on these results and well-established statistical theory, we advocate that investigators avoid rarefying altogether. We have provided microbiome-specific extensions to these tools in the R package, phyloseq.
22 pages, 5 figures, 2 supplementary sections
References in corpus (1)
Cited by in corpus (31)
- Sparse and compositionally robust inference of microbial ecological networks
- Bugs as Features (Part II): A Perspective on Enriching Microbiome-Gut-Brain Axis Analyses
- Bugs as Features (Part I): Concepts and Foundations for the Compositional Data Analysis of the Microbiome-Gut-Brain Axis
- Zero-inflated Poisson Factor Model with Application to Microbiome Absolute Abundance Data
- Bayesian Multinomial Logistic Normal Models through Marginally Latent Matrix-T Processes
- An ensemble approach to the structure-function problem in microbial communities
- Bayesian Nonparametric Ordination for the Analysis of Microbial Communities
- Estimation of symbiotic bacterial structure in a sustainable seagrass ecosystem on recycled management
- Quantitative Comparison of Abundance Structures of Generalized Communities: From B-Cell Receptor Repertoires to Microbiomes
- Elementary methods provide more replicable results in microbial differential abundance analysis
- Testing for differential abundance in compositional counts data, with application to microbiome studies
- Poisson PCA: Poisson Measurement Error corrected PCA, with Application to Microbiome Data
- A Power Analysis of the Conditional Randomization Test and Knockoffs
- COVID-19 Heterogeneity in Islands Chain Environment
- Adaptive gPCA: A method for structured dimensionality reduction
- The Block Bootstrap Method for Longitudinal Microbiome Data
- Bayesian Modeling of Microbiome Data for Differential Abundance Analysis
- Mixed Effect Dirichlet-Tree Multinomial for Longitudinal Microbiome Data and Weight Prediction
- Species richness estimation with high diversity but spurious singletons
- Conditional regression based on a multivariate zero-inflated logistic normal model for microbiome relative abundance data
- High-dimensional Log-Error-in-Variable Regression with Applications to Microbial Compositional Data Analysis
- A Common Atom Model for the Bayesian Nonparametric Analysis of Nested Data
- IFAA: Robust association identification and Inference For Absolute Abundance in microbiome analyses
- Metagenomic Analysis using Phylogenetic Placement -- A Review of the First Decade
- Inference for changes in biodiversity
- Metagenomics for clinical diagnostics: technologies and informatics
- Rough Set Microbiome Characterisation
- Inverse Probability Weighting-based Mediation Analysis for Microbiome Data
- Inference of Microbial Interactions Using Copula Models with Mixture Margins
- It's All Relative: New Regression Paradigm for Microbiome Compositional Data
- Phylogenetics and the human microbiome