Fast matrix-free evaluation of discontinuous Galerkin finite element operators
arXiv:1711.03590 · doi:10.1145/3325864
Abstract
We present an algorithmic framework for matrix-free evaluation of discontinuous Galerkin finite element operators based on sum factorization on quadrilateral and hexahedral meshes. We identify a set of kernels for fast quadrature on cells and faces targeting a wide class of weak forms originating from linear and nonlinear partial differential equations. Different algorithms and data structures for the implementation of operator evaluation are compared in an in-depth performance analysis. The sum factorization kernels are optimized by vectorization over several cells and faces and an even-odd decomposition of the one-dimensional compute kernels. In isolation our implementation then reaches up to 60\% of arithmetic peak on Intel Haswell and Broadwell processors and up to 50\% of arithmetic peak on Intel Knights Landing. The full operator evaluation reaches only about half that throughput due to memory bandwidth limitations from loading the input and output vectors, MPI ghost exchange, as well as handling variable coefficients and the geometry. Our performance analysis shows that the results are often within 10\% of the available memory bandwidth for the proposed implementation, with the exception of the Cartesian mesh case where the cost of gather operations and MPI communication are more substantial.
References in corpus (5)
- A performance comparison of continuous and discontinuous Galerkin methods with fast multigrid solvers
- A high-order semi-explicit discontinuous Galerkin solver for 3D incompressible flow with application to DNS and LES of turbulent channel flow
- Efficiency of high-performance discontinuous Galerkin spectral element methods for under-resolved turbulent incompressible flows
- A matrix-free high-order discontinuous Galerkin compressible Navier-Stokes solver: A performance comparison of compressible and incompressible formulations for turbulent incompressible flows
- Efficient Explicit Time Stepping of High Order Discontinuous Galerkin Schemes for Waves
Cited by in corpus (35)
- MFEM: a modular finite element methods library
- The deal.II finite element library: design, features, and insights
- Robust and efficient discontinuous Galerkin methods for under-resolved turbulent incompressible flows
- Efficiency of high-performance discontinuous Galerkin spectral element methods for under-resolved turbulent incompressible flows
- A Flexible, Parallel, Adaptive Geometric Multigrid method for FEM
- Hybrid multigrid methods for high-order discontinuous Galerkin discretizations
- Matrix-free multigrid solvers for phase-field fracture problems
- A matrix-free high-order discontinuous Galerkin compressible Navier-Stokes solver: A performance comparison of compressible and incompressible formulations for turbulent incompressible flows
- On the implementation of a robust and efficient finite element-based parallel solver for the compressible Navier-Stokes equations
- A matrix-free high-order solver for the numerical solution of cardiac electrophysiology
- High-order semi-Lagrangian kinetic scheme for compressible turbulence
- Scaling to the stars -- a linearly scaling elliptic solver for -multigrid
- High-order arbitrary Lagrangian-Eulerian discontinuous Galerkin methods for the incompressible Navier-Stokes equations
- Fast Tensor Product Schwarz Smoothers for High-Order Discontinuous Galerkin Methods
- Accelerating High-Order Mesh Optimization Using Finite Element Partial Assembly on GPUs
- End-to-end GPU acceleration of low-order-refined preconditioning for high-order finite element discretizations
- Conservative and accurate solution transfer between high-order and low-order refined finite element spaces
- A consistent diffuse-interface model for two-phase flow problems with rapid evaporation
- Matrix-Free Higher-Order Finite Element Methods for Hyperelasticity
- An implementation of tensor product patch smoothers on GPU
- Smoothers with localized residual computations for geometric multigrid methods
- A consistent diffuse-interface finite element approach to rapid melt--vapor dynamics with application to metal additive manufacturing
- A Generalized Probabilistic Learning Approach for Multi-Fidelity Uncertainty Propagation in Complex Physical Simulations
- COMMET: orders-of-magnitude speed-up in finite element method via batch-vectorized neural constitutive updates
- Fast hardware-aware matrix-free algorithm for higher-order finite-element discretized matrix multivector products on distributed systems
- Double-grid quadrature with interpolation-projection (DoGIP) as a novel discretisation approach: An application to FEM on simplexes
- A quantitative comparison of high-order asymptotic-preserving and asymptotically-accurate IMEX methods for the Euler equations with non-ideal gases
- Multilevel Interior Penalty Methods on GPUs
- An asynchronous discontinuous Galerkin method for massively parallel PDE solvers
- Scalability of High-Performance PDE Solvers
- Linearizing the hybridizable discontinuous Galerkin method: A linearly scaling operator
- A Hermite-like basis for faster matrix-free evaluation of interior penalty discontinuous Galerkin operators
- Tensor-product vertex patch smoothers for biharmonic problems
- Higher-Order Discontinuous Galerkin Splitting Schemes for Fluids with Variable Viscosity
- Improving the scalability of a high-order atmospheric dynamics solver based on the deal.II library