Scaling metagenome sequence assembly with probabilistic de Bruijn graphs
arXiv:1112.4193 · doi:10.1073/pnas.1121464109
Abstract
Deep sequencing has enabled the investigation of a wide range of environmental microbial ecosystems, but the high memory requirements for {\em de novo} assembly of short-read shotgun sequencing data from these complex populations are an increasingly large practical barrier. Here we introduce a memory-efficient graph representation with which we can analyze the k-mer connectivity of metagenomic samples. The graph representation is based on a probabilistic data structure, a Bloom filter, that allows us to efficiently store assembly graphs in as little as 4 bits per k-mer, albeit inexactly. We show that this data structure accurately represents DNA assembly graphs in low memory. We apply this data structure to the problem of partitioning assembly graphs into components as a prelude to assembly, and show that this reduces the overall memory requirements for {\em de novo} assembly of metagenomes. On one soil metagenome assembly, this approach achieves a nearly 40-fold decrease in the maximum memory requirements for assembly. This probabilistic graph representation is a significant theoretical advance in storing assembly graphs and also yields immediate leverage on metagenomic assembly.
Cited by in corpus (12)
- These are not the k-mers you are looking for: efficient online k-mer counting using a probabilistic data structure
- Survey and Taxonomy of Lossless Graph Compression and Space-Efficient Graph Representations
- Improving transcriptome assembly through error correction of high-throughput sequence reads
- Data structures to represent a set of k-long DNA sequences
- khmer: Working with Big Data in Bioinformatics
- On the representation of de Bruijn graphs
- Illumina Sequencing Artifacts Revealed by Connectivity Analysis of Metagenomic Datasets
- Assembling large, complex environmental metagenomes
- Compression of high throughput sequencing data with probabilistic de Bruijn graph
- Using cascading Bloom filters to improve the memory usage for de Brujin graphs
- Wavelet analysis on symbolic sequences and two-fold de Bruijn sequences
- Phylogenetics and the human microbiome