A Faster Grammar-Based Self-Index
arXiv:1109.3954
Abstract
To store and search genomic databases efficiently, researchers have recently started building compressed self-indexes based on grammars. In this paper we show how, given a straight-line program with rules for a string (S [1..n]) whose LZ77 parse consists of phrases, we can store a self-index for in $\Oh{r + z \log \log n}$ space such that, given a pattern (P [1..m]), we can list the $\occ$ occurrences of in in $\Oh{m^2 + \occ \log \log n}$ time. If the straight-line program is balanced and we accept a small probability of building a faulty index, then we can reduce the $\Oh{m^2}$ term to $\Oh{m \log m}$. All previous self-indexes are larger or slower in the worst case.
journal version of LATA '12 paper
Cited by in corpus (13)
- Linear Time Lempel-Ziv Factorization: Simple, Fast, Small
- Universal Compressed Text Indexing
- Lightweight Lempel-Ziv Parsing
- Time-Space Trade-Offs for Lempel-Ziv Compressed Indexing
- Universal Indexes for Highly Repetitive Document Collections
- Relative Suffix Trees
- Linear-Space Substring Range Counting over Polylogarithmic Alphabets
- Online Self-Indexed Grammar Compression
- AliBI: An Alignment-Based Index for Genomic Datasets
- Grammar Index By Induced Suffix Sorting
- Heaviest Induced Ancestors and Longest Common Substrings
- Orthogonal Range Searching for Text Indexing
- On Two LZ78-style Grammars: Compression Bounds and Compressed-Space Computation