Consistent and Flexible Selectivity Estimation for High-Dimensional Data
arXiv:2005.09908 · doi:10.1145/3448016.3452772
Abstract
Selectivity estimation aims at estimating the number of database objects that satisfy a selection criterion. Answering this problem accurately and efficiently is essential to many applications, such as density estimation, outlier detection, query optimization, and data integration. The estimation problem is especially challenging for large-scale high-dimensional data due to the curse of dimensionality, the large variance of selectivity across different queries, and the need to make the estimator consistent (i.e., the selectivity is non-decreasing in the threshold). We propose a new deep learning-based model that learns a query-dependent piecewise linear function as selectivity estimator, which is flexible to fit the selectivity curve of any distance function and query object, while guaranteeing that the output is non-decreasing in the threshold. To improve the accuracy for large datasets, we propose to partition the dataset into multiple disjoint subsets and build a local model on each of them. We perform experiments on real datasets and show that the proposed model consistently outperforms state-of-the-art models in accuracy in an efficient way and is useful for real applications.
Published at ACM SIGMOD Conference 2021
References in corpus (10)
- Neo: A Learned Query Optimizer
- Deep Entity Matching with Pre-Trained Language Models
- Deep Unsupervised Cardinality Estimation
- QuickSel: Quick Selectivity Learning with Mixture Models
- Deep Lattice Networks and Partial Monotonic Functions
- RadixSpline: A Single-Pass Learned Index
- An Empirical Analysis of Deep Learning for Cardinality Estimation
- Monotonic Cardinality Estimation of Similarity Selection: A Deep Learning Approach
- Learning to Sample: Counting with Complex Queries
- Consistent and Flexible Selectivity Estimation for High-Dimensional Data