Chip-level and multi-node analysis of energy-optimized lattice-Boltzmann CFD simulations
arXiv:1304.7664 · doi:10.1002/cpe.3489
Abstract
Memory-bound algorithms show complex performance and energy consumption behavior on multicore processors. We choose the lattice-Boltzmann method (LBM) on an Intel Sandy Bridge cluster as a prototype scenario to investigate if and how single-chip performance and power characteristics can be generalized to the highly parallel case. First we perform an analysis of a sparse-lattice LBM implementation for complex geometries. Using a single-core performance model, we predict the intra-chip saturation characteristics and the optimal operating point in terms of energy to solution as a function of implementation details, clock frequency, vectorization, and number of active cores per chip. We show that high single-core performance and a correct choice of the number of active cores per chip are the essential optimizations for lowest energy to solution at minimal performance degradation. Then we extrapolate to the MPI-parallel level and quantify the energy-saving potential of various optimizations and execution modes, where we find these guidelines to be even more important, especially when communication overhead is non-negligible. In our setup we could achieve energy savings of 35% in this case, compared to a naive approach. We also demonstrate that a simple non-reflective reduction of the clock speed leaves most of the energy saving potential unused.
23 pages, 13 figures; post-peer-review version
References in corpus (4)
- Comparison of different Propagation Steps for the Lattice Boltzmann Method
- Exploring performance and power properties of modern multicore chips via simple machine models
- Pushing the limits for medical image reconstruction on recent standard multicore processors
- Introducing a Performance Model for Bandwidth-Limited Loop Kernels
Cited by in corpus (17)
- The Ecological Impact of High-performance Computing in Astrophysics
- Evaluation of DVFS techniques on modern HPC processors and accelerators for energy-aware applications
- Oort cloud Ecology II: Extra-solar Oort clouds and the origin of asteroidal interlopers
- Kerncraft: A Tool for Analytic Performance Modeling of Loop Kernels
- Lattice Boltzmann Benchmark Kernels as a Testbed for Performance Analysis
- Propagation and Decay of Injected One-Off Delays on Clusters: A Case Study
- On the accuracy and usefulness of analytic energy models for contemporary multicore processors
- Modeling and analyzing performance for highly optimized propagation steps of the lattice Boltzmann method on sparse lattices
- Performance analysis of the Kahan-enhanced scalar product on current multi- and manycore processors
- SPEChpc 2021 Benchmarks on Ice Lake and Sapphire Rapids Infiniband Clusters: A Performance and Energy Case Study
- Bridging the Architecture Gap: Abstracting Performance-Relevant Properties of Modern Server Processors
- Oort Cloud Ecology. III. The Sun left the parent star cluster shortly after the giant planets formed
- Energy efficiency: a Lattice Boltzmann study
- The Steady State of Intermediate-Mass Black Holes Near a Supermassive Black Hole
- Extreme Scale-out SuperMUC Phase 2 - lessons learned
- The origin and evolution of wide Jupiter Mass Binary Objects in young stellar clusters
- Edge Intelligence for Energy-efficient Computation Offloading and Resource Allocation in 5G Beyond