Multiseed Lossless Filtration
arXiv:0901.3215 · doi:10.1109/TCBB.2005.12
Abstract
We study a method of seed-based lossless filtration for approximate string matching and related bioinformatics applications. The method is based on a simultaneous use of several spaced seeds rather than a single seed as studied by Burkhardt and Kärkkäinen [1]. We present algorithms to compute several important parameters of seed families, study their combinatorial properties, and describe several techniques to construct efficient families. We also report a large-scale application of the proposed technique to the problem of oligonucleotide selection for an EST sequence database.
References in corpus (1)
Cited by in corpus (7)
- A unifying framework for seed sensitivity and its application to subset seeds
- A Coverage Criterion for Spaced Seeds and its Applications to Support Vector Machine String Kernels and k-Mer Distances
- A unifying framework for seed sensitivity and its application to subset seeds (Extended abstract)
- Progressive Mauve: Multiple alignment of genomes with gene flux and rearrangement
- Seed design framework for mapping SOLiD reads
- Languages of lossless seeds
- Approximate String Matching using a Bidirectional Index