A unifying framework for seed sensitivity and its application to subset seeds
arXiv:cs/0601116 · doi:10.1142/S0219720006001977
Abstract
We propose a general approach to compute the seed sensitivity, that can be applied to different definitions of seeds. It treats separately three components of the seed sensitivity problem -- a set of target alignments, an associated probability distribution, and a seed model -- that are specified by distinct finite automata. The approach is then applied to a new concept of subset seeds for which we propose an efficient automaton construction. Experimental results confirm that sensitive subset seeds can be efficiently designed using our approach, and can then be used in similarity search producing better results than ordinary spaced seeds.
References in corpus (1)
Cited by in corpus (6)
- Spaced seeds improve k-mer-based metagenomic classification
- RasBhari: optimizing spaced seeds for database searching, read mapping and alignment-free sequence comparison
- On subset seeds for protein alignment
- A Coverage Criterion for Spaced Seeds and its Applications to Support Vector Machine String Kernels and k-Mer Distances
- Extraction of long k-mers using spaced seeds
- Seed design framework for mapping SOLiD reads