Petascale turbulence simulation using a highly parallel fast multipole method on GPUs
arXiv:1106.5273 · doi:10.1016/j.cpc.2012.09.011
Abstract
This paper reports large-scale direct numerical simulations of homogeneous-isotropic fluid turbulence, achieving sustained performance of 1.08 petaflop/s on gpu hardware using single precision. The simulations use a vortex particle method to solve the Navier-Stokes equations, with a highly parallel fast multipole method (FMM) as numerical engine, and match the current record in mesh size for this application, a cube of 4096^3 computational points solved with a spectral method. The standard numerical approach used in this field is the pseudo-spectral method, relying on the FFT algorithm as numerical engine. The particle-based simulations presented in this paper quantitatively match the kinetic energy spectrum obtained with a pseudo-spectral method, using a trusted code. In terms of parallel performance, weak scaling results show the fmm-based vortex method achieving 74% parallel efficiency on 4096 processes (one gpu per mpi process, 3 gpus per node of the TSUBAME-2.0 system). The FFT-based spectral method is able to achieve just 14% parallel efficiency on the same number of mpi processes (using only cpu cores), due to the all-to-all communication pattern of the FFT algorithm. The calculation time for one time step was 108 seconds for the vortex method and 154 seconds for the spectral method, under these conditions. Computing with 69 billion particles, this work exceeds by an order of magnitude the largest vortex method calculations to date.
References in corpus (3)
Cited by in corpus (22)
- A Tuned and Scalable Fast Multipole Method as a Preeminent Algorithm for Exascale Systems
- Reviving the Vortex Particle Method: A Stable Formulation for Meshless Large Eddy Simulation
- FMM-based vortex method for simulation of isotropic turbulence on GPUs, compared with a spectral method
- Computational Physics on Graphics Processing Units
- Flexibly imposing periodicity in kernel independent FMM: A Multipole-To-Local operator approach
- From Piz Daint to the Stars: Simulation of Stellar Mergers using High-Level Abstractions
- Canonical symplectic structure and structure-preserving geometric algorithms for Schrödinger-Maxwell systems
- A method to compute periodic sums
- Fast Multipole Method as a Matrix-Free Hierarchical Low-Rank Approximation
- Fast electrostatic solvers for kinetic Monte Carlo simulations
- RPYFMM: Parallel Adaptive Fast Multipole Method for Rotne-Prager-Yamakawa Tensor in Biomolecular Hydrodynamics Simulations
- An FMM Based on Dual Tree Traversal for Many-core Architectures
- A Smooth Partition of Unity Finite Element Method for Vortex Particle Regularization
- Accelerating Brain Simulations with the Fast Multipole Method
- A Particle Method without Remeshing
- On Particles and Splines in Bounded Domains
- GPU parallelization of a hybrid pseudospectral fluid turbulence framework using CUDA
- Multiscale Computing in the Exascale Era
- Fast multipole networks
- Tracking the vortex motion by using Brownian fluid particles
- Distributed and Adaptive Fast Multipole Method In Three Dimensions
- A GPU-Parallelized Interpolation-Based Fast Multipole Method for the Relativistic Space-Charge Field Calculation