Communication Steps for Parallel Query Processing
arXiv:1306.5972
Abstract
We consider the problem of computing a relational query on a large input database of size , using a large number of servers. The computation is performed in rounds, and each server can receive only bits of data, where is a parameter that controls replication. We examine how many global communication steps are needed to compute . We establish both lower and upper bounds, in two settings. For a single round of communication, we give lower bounds in the strongest possible model, where arbitrary bits may be exchanged; we show that any algorithm requires , where is the fractional vertex cover of the hypergraph of . We also give an algorithm that matches the lower bound for a specific class of databases. For multiple rounds of communication, we present lower bounds in a model where routing decisions for a tuple are tuple-based. We show that for the class of tree-like queries there exists a tradeoff between the number of rounds and the space exponent . The lower bounds for multiple rounds are the first of their kind. Our results also imply that transitive closure cannot be computed in O(1) rounds of communication.
References in corpus (1)
Cited by in corpus (16)
- Worst-Case Optimal Algorithms for Parallel Query Processing
- Research Directions for Principles of Data Management (Dagstuhl Perspectives Workshop 16151)
- EmptyHeaded: A Relational Engine for Graph Processing
- Skew in Parallel Query Processing
- Parallel Algorithms for Geometric Graph Problems
- MapReduce Meets Fine-Grained Complexity: MapReduce Algorithms for APSP, Matrix Multiplication, 3-SUM, and Beyond
- Tribes Is Hard in the Message Passing Model
- Communication Cost in Parallel Query Processing
- Parallel Evaluation of Multi-Semi-Joins
- On the Computational Complexity of MapReduce
- Parallel-Correctness and Containment for Conjunctive Queries with Union and Negation
- Parallel-Correctness and Transferability for Conjunctive Queries
- Parallel Batch-Dynamic Graphs: Algorithms and Lower Bounds
- Going for Speed: Sublinear Algorithms for Dense r-CSPs
- Sparse Hopsets in Congested Clique
- Distributed Statistical Estimation of Matrix Products with Applications