Hercules Against Data Series Similarity Search
arXiv:2212.13297 · doi:10.14778/3547305.3547308
Abstract
We propose Hercules, a parallel tree-based technique for exact similarity search on massive disk-based data series collections. We present novel index construction and query answering algorithms that leverage different summarization techniques, carefully schedule costly operations, optimize memory and disk accesses, and exploit the multi-threading and SIMD capabilities of modern hardware to perform CPU-intensive calculations. We demonstrate the superiority and robustness of Hercules with an extensive experimental evaluation against state-of-the-art techniques, using many synthetic and real datasets, and query workloads of varying difficulty. The results show that Hercules performs up to one order of magnitude faster than the best competitor (which is not always the same). Moreover, Hercules is the only index that outperforms the optimized scan on all scenarios, including the hard query workloads on disk-based datasets. This paper was published in the Proceedings of the VLDB Endowment, Volume 15, Number 10, June 2022.
References in corpus (3)
Cited by in corpus (8)
- dCAM: Dimension-wise Class Activation Map for Explaining Multivariate Data Series Classification
- DET-LSH: A Locality-Sensitive Hashing Scheme with Dynamic Encoding Tree for Approximate Nearest Neighbor Search
- Dumpy: A Compact and Adaptive Index for Large Data Series Collections
- SEAnet: A Deep Learning Architecture for Data Series Similarity Search
- LIRA: A Learning-based Query-aware Partition Framework for Large-scale ANN Search
- LeaFi: Data Series Indexes on Steroids with Learned Filters
- Raising the ClaSS of Streaming Time Series Segmentation
- HetFS: A Method for Fast Similarity Search with Ad-hoc Meta-paths on Heterogeneous Information Networks