Nodal Discontinuous Galerkin Methods on Graphics Processors
arXiv:0901.1024 · doi:10.1016/j.jcp.2009.06.041
Abstract
Discontinuous Galerkin (DG) methods for the numerical solution of partial differential equations have enjoyed considerable success because they are both flexible and robust: They allow arbitrary unstructured geometries and easy control of accuracy without compromising simulation stability. Lately, another property of DG has been growing in importance: The majority of a DG operator is applied in an element-local way, with weak penalty-based element-to-element coupling. The resulting locality in memory access is one of the factors that enables DG to run on off-the-shelf, massively parallel graphics processors (GPUs). In addition, DG's high-order nature lets it require fewer data points per represented wavelength and hence fewer memory accesses, in exchange for higher arithmetic intensity. Both of these factors work significantly in favor of a GPU implementation of DG. Using a single US$400 Nvidia GTX 280 GPU, we accelerate a solver for Maxwell's equations on a general 3D unstructured grid by a factor of 40 to 60 relative to a serial computation on a current-generation CPU. In many cases, our algorithms exhibit full use of the device's available memory bandwidth. Example computations achieve and surpass 200 gigaflops/s of net application-level floating point work. In this article, we describe and derive the techniques used to reach this level of performance. In addition, we present comprehensive data on the accuracy and runtime behavior of the method.
33 pages, 12 figures, 4 tables
Cited by in corpus (58)
- PyCUDA and PyOpenCL: A Scripting-Based Approach to GPU Run-Time Code Generation
- On discretely entropy conservative and entropy stable discontinuous Galerkin methods
- Fast matrix-free evaluation of discontinuous Galerkin finite element operators
- Viscous Shock Capturing in a Time-Explicit Discontinuous Galerkin Method
- OCCA: A unified approach to multi-threading languages
- GPU-accelerated discontinuous Galerkin methods on hybrid meshes
- GPU Accelerated Discontinuous Galerkin Methods for Shallow Water Equations
- Polymer Field-Theory Simulations on Graphics Processing Units
- GPU performance analysis of a nodal discontinuous Galerkin method for acoustic and elastic models
- An entropy stable discontinuous Galerkin method for the shallow water equations on curvilinear meshes with wet/dry fronts accelerated by GPUs
- Discontinuous Galerkin method for the spherically reduced BSSN system with second-order operators
- Numerical integration on GPUs for higher order finite elements
- Nodal Discontinuous Galerkin Simulations for Reverse-Time Migration on GPU Clusters
- Discontinuous Galerkin methods on graphics processing units for nonlinear hyperbolic conservation laws
- Finite Element Integration on GPUs
- Finite element numerical integration for first order approximations on multi-core architectures
- Efficient low-order refined preconditioners for high-order matrix-free continuous and discontinuous Galerkin methods
- A Discontinuous Galerkin Time-Domain Method with Dynamically Adaptive Cartesian Meshes for Computational Electromagnetics
- Vectorized OpenCL implementation of numerical integration for higher order finite elements
- A New Sparse Matrix Vector Multiplication GPU Algorithm Designed for Finite Element Problems
- A weight-adjusted discontinuous Galerkin method for the poroelastic wave equation: penalty fluxes and micro-heterogeneities
- An Energy Based Discontinuous Galerkin Method for Coupled Elasto-Acoustic Wave Equations in Second Order Form
- High-order matrix-free incompressible flow solvers with GPU acceleration and low-order refined preconditioners
- Realizability-Preserving DG-IMEX Method for the Two-Moment Model of Fermion Transport
- ForestClaw: Hybrid forest-of-octrees AMR for hyperbolic conservation laws
- Weight-adjusted discontinuous Galerkin methods: matrix-valued weights and elastic wave propagation in heterogeneous media
- A variable high-order shock-capturing finite difference method with GP-WENO
- Scaling to the stars -- a linearly scaling elliptic solver for -multigrid
- Bound-Preserving Discontinuous Galerkin Methods for Conservative Phase Space Advection in Curvilinear Coordinates
- A DG-IMEX method for two-moment neutrino transport: Nonlinear solvers for neutrino-matter coupling
- Enclave Tasking for Discontinuous Galerkin Methods on Dynamically Adaptive Meshes
- Efficient Entropy-Stable Discontinuous Spectral-Element Methods Using Tensor-Product Summation-by-Parts Operators on Triangles and Tetrahedra
- Loo.py: From Fortran to performance via transformation and substitution rules
- A Discontinuous Galerkin Finite Element Model for Compound Flood Simulations
- A curved-element unstructured discontinuous Galerkin method on GPUs for the Euler equations
- Discovering Artificial Viscosity Models for Discontinuous Galerkin Approximation of Conservation Laws using Physics-Informed Machine Learning
- High order weight-adjusted discontinuous Galerkin methods for wave propagation on moving curved meshes
- An implementation of tensor product patch smoothers on GPU
- Weight-adjusted discontinuous Galerkin methods: wave propagation in heterogeneous media
- GPU-accelerated Bernstein-Bezier discontinuous Galerkin methods for wave problems
- A hybridizable discontinuous Galerkin method with characteristic variables for Helmholtz problems
- Weight-adjusted discontinuous Galerkin methods: curvilinear meshes
- Efficient discontinuous Galerkin finite element methods via Bernstein polynomials
- A discontinuous Galerkin fast spectral method for multi-species full Boltzmann on streaming multi-processors
- Complete PISO and SIMPLE solvers on Graphics Processing Units
- Orthogonal bases for vertex-mapped pyramids
- Multilevel Interior Penalty Methods on GPUs
- Bernstein-Bezier weight-adjusted discontinuous Galerkin methods for wave propagation in heterogeneous media
- A mechanism for balancing accuracy and scope in cross-machine black-box GPU performance modeling
- High-Order Discontinuous Galerkin Methods by GPU Metaprogramming
- Parallel kinetic schemes for conservation laws, with large time steps
- GPU Accelerated Finite Element Assembly with Runtime Compilation
- Reduced storage nodal discontinuous Galerkin methods on semi-structured prismatic meshes
- GPU-accelerated discontinuous Galerkin methods on polytopic meshes
- GPU-based parallel simulations of the Gatenby-Gawlinski model with anisotropic, heterogeneous acid diffusion
- Edge coloring in unstructured CFD codes
- Solving Wave Equations on Unstructured Geometries
- Heterogeneous Computing on Mixed Unstructured Grids with PyFR