Graph Expansion and Communication Costs of Fast Matrix Multiplication
arXiv:1109.1693 · doi:10.1145/1989493.1989495
Abstract
The communication cost of algorithms (also known as I/O-complexity) is shown to be closely related to the expansion properties of the corresponding computation graphs. We demonstrate this on Strassen's and other fast matrix multiplication algorithms, and obtain first lower bounds on their communication costs. In the sequential case, where the processor has a fast memory of size , too small to store three -by- matrices, the lower bound on the number of words moved between fast and slow memory is, for many of the matrix multiplication algorithms, , where is the exponent in the arithmetic count (e.g., for Strassen, and for conventional matrix multiplication). With parallel processors, each with fast memory of size , the lower bound is times smaller. These bounds are attainable both for sequential and for parallel algorithms and hence optimal. These bounds can also be attained by many fast algorithms in linear algebra (e.g., algorithms for LU, QR, and solving the Sylvester equation).
References in corpus (2)
Cited by in corpus (6)
- Graph Expansion and Communication Costs of Fast Matrix Multiplication
- Strong Scaling of Matrix Multiplication Algorithms and Memory-Independent Communication Lower Bounds
- Communication-Optimal Parallel Algorithm for Strassen's Matrix Multiplication
- Tight Bounds for Low Dimensional Star Stencils in the Parallel External Memory Model
- Improving the numerical stability of fast matrix multiplication
- Fast Strassen-based Parallel Multiplication