KMC 2: Fast and resource-frugal -mer counting
arXiv:1407.1507 · doi:10.1093/bioinformatics/btv022
Abstract
Motivation: Building the histogram of occurrences of every -symbol long substring of nucleotide data is a standard step in many bioinformatics applications, known under the name of -mer counting. Its applications include developing de Bruijn graph genome assemblers, fast multiple sequence alignment and repeat detection. The tremendous amounts of NGS data require fast algorithms for -mer counting, preferably using moderate amounts of memory. Results: We present a novel method for -mer counting, on large datasets at least twice faster than the strongest competitors (Jellyfish~2, KMC~1), using about 12\,GB (or less) of RAM memory. Our disk-based method bears some resemblance to MSPKmerCounter, yet replacing the original minimizers with signatures (a carefully selected subset of all minimizers) and using -mers allows to significantly reduce the I/O, and a highly parallel overall architecture allows to achieve unprecedented processing speeds. For example, KMC~2 allows to count the 28-mers of a human reads collection with 44-fold coverage (106\,GB of compressed size) in about 20 minutes, on a 6-core Intel i7 PC with an SSD. Availability: KMC~2 is freely available at http://sun.aei.polsl.pl/kmc. Contact: [email protected]
Cited by in corpus (12)
- Internal Pattern Matching Queries in a Text and Applications
- RECKONER: Read Error Corrector Based on KMC
- Multiple Comparative Metagenomics using Multiset k-mer Counting
- Correcting Illumina sequencing errors for human data
- Compression of high throughput sequencing data with probabilistic de Bruijn graph
- Alignment-Free Sequence Analysis and Applications
- Gerbil: A Fast and Memory-Efficient -mer Counter with GPU-Support
- A bloated FM-index reducing the number of cache misses during the search
- Analyzing Big Datasets of Genomic Sequences: Fast and Scalable Collection of k-mer Statistics
- Sorting Data on Ultra-Large Scale with RADULS. New Incarnation of Radix Sort
- Even better correction of genome sequencing data
- Disk-based genome sequencing data compression