Skew Strikes Back: New Developments in the Theory of Join Algorithms
arXiv:1310.3314
Abstract
Evaluating the relational join is one of the central algorithmic and most well-studied problems in database systems. A staggering number of variants have been considered including Block-Nested loop join, Hash-Join, Grace, Sort-merge for discussions of more modern issues). Commercial database engines use finely tuned join heuristics that take into account a wide variety of factors including the selectivity of various predicates, memory, IO, etc. In spite of this study of join queries, the textbook description of join processing is suboptimal. This survey describes recent results on join algorithms that have provable worst-case optimality runtime guarantees. We survey recent work and provide a simpler and unified description of these algorithms that we hope is useful for theory-minded readers, algorithm designers, and systems implementors.
Cited by in corpus (10)
- Ringo: Interactive Graph Analytics on Big-Memory Machines
- Flexible Caching in Trie Joins
- EmptyHeaded: A Relational Engine for Graph Processing
- AC/DC: In-Database Learning Thunderstruck
- FAQ: Questions Asked Frequently
- Join Processing for Graph Patterns: An Old Dog with New Tricks
- Joins via Geometric Resolutions: Worst-case and Beyond
- A Near-Optimal Parallel Algorithm for Joining Binary Relations
- Fast Distributed Complex Join Processing
- Juggling Functions Inside a Database