Low Rank Learning for Offline Query Optimization
arXiv:2504.06399 · doi:10.1145/3725412
Abstract
Recent deployments of learned query optimizers use expensive neural networks and ad-hoc search policies. To address these issues, we introduce \textsc{LimeQO}, a framework for offline query optimization leveraging low-rank learning to efficiently explore alternative query plans with minimal resource usage. By modeling the workload as a partially observed, low-rank matrix, we predict unobserved query plan latencies using purely linear methods, significantly reducing computational overhead compared to neural networks. We formalize offline exploration as an active learning problem, and present simple heuristics that reduces a 3-hour workload to 1.5 hours after just 1.5 hours of exploration. Additionally, we propose a transductive Tree Convolutional Neural Network (TCNN) that, despite higher computational costs, achieves the same workload reduction with only 0.5 hours of exploration. Unlike previous approaches that place expensive neural networks directly in the query processing ``hot'' path, our approach offers a low-overhead solution and a no-regressions guarantee, all without making assumptions about the underlying DBMS. The code is available in \href{https://github.com/zixy17/LimeQO}{https://github.com/zixy17/LimeQO}.
To appear in SIGMOD 2025
References in corpus (12)
- Deep Learning based Recommender System: A Survey and New Perspectives
- Neo: A Learned Query Optimizer
- Bao: Learning to Steer Query Optimizers
- Deep Unsupervised Cardinality Estimation
- Apache Calcite: A Foundational Framework for Optimized Query Processing Over Heterogeneous Data Sources
- Balsa: Learning a Query Optimizer Without Expert Demonstrations
- ALECE: An Attention-based Learned Cardinality Estimator for SPJ Queries on Dynamic Workloads (Extended)
- Deploying a Steered Query Optimizer in Production at Microsoft
- Matrix completion with queries
- Robust Plan Evaluation based on Approximate Probabilistic Machine Learning
- Grid-AR: A Grid-based Booster for Learned Cardinality Estimation and Range Joins
- Falcon: Fair Active Learning using Multi-armed Bandits