Spaced seeds improve k-mer-based metagenomic classification
arXiv:1502.06256 · doi:10.1093/bioinformatics/btv419
Abstract
Metagenomics is a powerful approach to study genetic content of environmental samples that has been strongly promoted by NGS technologies. To cope with massive data involved in modern metagenomic projects, recent tools [4, 39] rely on the analysis of k-mers shared between the read to be classified and sampled reference genomes. Within this general framework, we show in this work that spaced seeds provide a significant improvement of classification accuracy as opposed to traditional contiguous k-mers. We support this thesis through a series a different computational experiments, including simulations of large-scale metagenomic projects. Scripts and programs used in this study, as well as supplementary material, are available from http://github.com/gregorykucherov/spaced-seeds-for-metagenomics.
23 pages
References in corpus (1)
Cited by in corpus (5)
- RasBhari: optimizing spaced seeds for database searching, read mapping and alignment-free sequence comparison
- Multiple Comparative Metagenomics using Multiset k-mer Counting
- Low-density locality-sensitive hashing boosts metagenomic binning
- Fast Approximation of Frequent -mers and Applications to Metagenomics
- Probabilistic Models of k-mer Frequencies (Extended Abstract)