Quantifying performance bottlenecks of stencil computations using the Execution-Cache-Memory model
arXiv:1410.5010 · doi:10.1145/2751205.2751240
Abstract
Stencil algorithms on regular lattices appear in many fields of computational science, and much effort has been put into optimized implementations. Such activities are usually not guided by performance models that provide estimates of expected speedup. Understanding the performance properties and bottlenecks by performance modeling enables a clear view on promising optimization opportunities. In this work we refine the recently developed Execution-Cache-Memory (ECM) model and use it to quantify the performance bottlenecks of stencil algorithms on a contemporary Intel processor. This includes applying the model to arrive at single-core performance and scalability predictions for typical corner case stencil loop kernels. Guided by the ECM model we accurately quantify the significance of "layer conditions," which are required to estimate the data traffic through the memory hierarchy, and study the impact of typical optimization approaches such as spatial blocking, strength reduction, and temporal blocking for their expected benefits. We also compare the ECM model to the widely known Roofline model.
10 pages, 8 figures. Added Roofline comparison and other minor improvements
Cited by in corpus (18)
- Automated Instruction Stream Throughput Prediction for Intel and AMD Microarchitectures
- Kerncraft: A Tool for Analytic Performance Modeling of Loop Kernels
- uiCA: Accurate Throughput Prediction of Basic Blocks on Recent Intel Microarchitectures
- ECM modeling and performance tuning of SpMV and Lattice QCD on A64FX
- Performance Modeling of Streaming Kernels and Sparse Matrix-Vector Multiplication on A64FX
- Automatic Loop Kernel Analysis and Performance Modeling With Kerncraft
- Automatic Throughput and Critical Path Analysis of x86 and ARM Assembly Kernels
- Propagation and Decay of Injected One-Off Delays on Clusters: A Case Study
- Performance prediction of finite-difference solvers for different computer architectures
- On the accuracy and usefulness of analytic energy models for contemporary multicore processors
- Revisiting Temporal Blocking Stencil Optimizations
- Desynchronization and Wave Pattern Formation in MPI-Parallel and Hybrid Memory-Bound Programs
- Performance of the low-rank tensor-train SVD (TT-SVD) for large dense tensors on modern multi-core CPUs
- Performance analysis of the Kahan-enhanced scalar product on current multicore processors
- Analytic Performance Modeling and Analysis of Detailed Neuron Simulations
- Performance analysis of the Kahan-enhanced scalar product on current multi- and manycore processors
- A domain-specific language and matrix-free stencil code for investigating electronic properties of Dirac and topological materials
- Model-Based Performance Analysis of the HyTeG Finite Element Framework