Pushing Memory Bandwidth Limitations Through Efficient Implementations of Block-Krylov Space Solvers on GPUs
arXiv:1710.09745 · doi:10.1016/j.cpc.2018.06.019
Abstract
Lattice quantum chromodynamics simulations in nuclear physics have benefited from a tremendous number of algorithmic advances such as multigrid and eigenvector deflation. These improve the time to solution but do not alleviate the intrinsic memory-bandwidth constraints of the matrix-vector operation dominating iterative solvers. Batching this operation for multiple vectors and exploiting cache and register blocking can yield a super-linear speed up. Block-Krylov solvers can naturally take advantage of such batched matrix-vector operations, further reducing the iterations to solution by sharing the Krylov space between solves. However, practical implementations typically suffer from the quadratic scaling in the number of vector-vector operations. Using the QUDA library, we present an implementation of a block-CG solver on NVIDIA GPUs which reduces the memory-bandwidth complexity of vector-vector operations from quadratic to linear. We present results for the HISQ discretization, showing a 5x speedup compared to highly-optimized independent Krylov solves on NVIDIA's SaturnV cluster.
15 pages, 14 figures, in press
References in corpus (7)
- Adaptive multigrid algorithm for the lattice Wilson-Dirac operator
- Lattice QCD as a video game
- Local coherence and deflation of the low quark modes in lattice QCD
- Adaptive Multigrid Algorithm for Lattice QCD
- A Framework for Lattice QCD Calculations on GPUs
- Conjugate gradient solvers on Intel Xeon Phi and NVIDIA GPUs
- Deflation Methods in Fermion Inverters
Cited by in corpus (7)
- Status and Future Perspectives for Lattice Gauge Theory Calculations to the Exascale and Beyond
- Multigrid for Staggered Lattice Fermions
- XAMG: A library for solving linear systems with multiple right-hand side vectors
- Towards Lattice Quantum Chromodynamics on FPGA devices
- Maximizing the Bang Per Bit
- RHMC with Block Solvers and Multiple Pseudofermions
- Optimizing Staggered Multigrid for Exascale performance