PetFMM--A dynamically load-balancing parallel fast multipole library
arXiv:0905.2637 · doi:10.1002/nme.2972
Abstract
Fast algorithms for the computation of -body problems can be broadly classified into mesh-based interpolation methods, and hierarchical or multiresolution methods. To this last class belongs the well-known fast multipole method (FMM), which offers O(N) complexity. This paper presents an extensible parallel library for -body interactions utilizing the FMM algorithm, built on the framework of PETSc. A prominent feature of this library is that it is designed to be extensible, with a view to unifying efforts involving many algorithms based on the same principles as the FMM and enabling easy development of scientific application codes. The paper also details an exhaustive model for the computation of tree-based -body algorithms in parallel, including both work estimates and communications estimates. With this model, we are able to implement a method to provide automatic, a priori load balancing of the parallel execution, achieving optimal distribution of the computational work among processors and minimal inter-processor communications. Using a client application that performs the calculation of velocity induced by vortex particles, ample verification and testing of the library was performed. Strong scaling results are presented with close to a million particles in up to 64 processors, including both speedup and parallel efficiency. The library is currently able to achieve over 85% parallel efficiency for 64 processors. The software library is open source under the PETSc license; this guarantees the maximum impact to the scientific community and encourages peer-based collaboration for the extensions and applications.
28 pages, 9 figures
Cited by in corpus (13)
- Biomolecular electrostatics using a fast multipole BEM on up to 512 GPUs and a billion unknowns
- A Tuned and Scalable Fast Multipole Method as a Preeminent Algorithm for Exascale Systems
- Optimal, scalable forward models for computing gravity anomalies
- Adaptive fast multipole methods on the GPU
- Fast Multipole Method for Gravitational Lensing. Application to High Magnification Quasar Microlensing
- How to obtain efficient GPU kernels: an illustration using FMM & FGT algorithms
- Dynamic autotuning of adaptive fast multipole methods on hybrid multicore CPU & GPU systems
- A parallel fast multipole method for a space-time boundary element method for the heat equation
- Pipelining the Fast Multipole Method over a Runtime System
- Real-space quadrature: a convenient, efficient representation for multipole expansions
- Extreme Scale FMM-Accelerated Boundary Integral Equation Solver for Wave Scattering
- Removing the Barrier to Scalability in Parallel FMM
- Distributed and Adaptive Fast Multipole Method In Three Dimensions