Kerncraft: A Tool for Analytic Performance Modeling of Loop Kernels
arXiv:1702.04653 · doi:10.1007/978-3-319-56702-0_1
Abstract
Achieving optimal program performance requires deep insight into the interaction between hardware and software. For software developers without an in-depth background in computer architecture, understanding and fully utilizing modern architectures is close to impossible. Analytic loop performance modeling is a useful way to understand the relevant bottlenecks of code execution based on simple machine models. The Roofline Model and the Execution-Cache-Memory (ECM) model are proven approaches to performance modeling of loop nests. In comparison to the Roofline model, the ECM model can also describes the single-core performance and saturation behavior on a multicore chip. We give an introduction to the Roofline and ECM models, and to stencil performance modeling using layer conditions (LC). We then present Kerncraft, a tool that can automatically construct Roofline and ECM models for loop nests by performing the required code, data transfer, and LC analysis. The layer condition analysis allows to predict optimal spatial blocking factors for loop nests. Together with the models it enables an ab-initio estimate of the potential benefits of loop blocking optimizations and of useful block sizes. In cases where LC analysis is not easily possible, Kerncraft supports a cache simulator as a fallback option. Using a 25-point long-range stencil we demonstrate the usefulness and predictive power of the Kerncraft tool.
22 pages, 5 figures
References in corpus (4)
- Quantifying performance bottlenecks of stencil computations using the Execution-Cache-Memory model
- GHOST: Building blocks for high performance sparse linear algebra on heterogeneous systems
- Automatic Loop Kernel Analysis and Performance Modeling With Kerncraft
- Performance analysis of the Kahan-enhanced scalar product on current multi- and manycore processors
Cited by in corpus (10)
- Automated Instruction Stream Throughput Prediction for Intel and AMD Microarchitectures
- uiCA: Accurate Throughput Prediction of Basic Blocks on Recent Intel Microarchitectures
- ECM modeling and performance tuning of SpMV and Lattice QCD on A64FX
- Performance Modeling of Streaming Kernels and Sparse Matrix-Vector Multiplication on A64FX
- Automatic Throughput and Critical Path Analysis of x86 and ARM Assembly Kernels
- On the accuracy and usefulness of analytic energy models for contemporary multicore processors
- Analytic Performance Modeling and Analysis of Detailed Neuron Simulations
- Boosting Performance Optimization with Interactive Data Movement Visualization
- A mechanism for balancing accuracy and scope in cross-machine black-box GPU performance modeling
- Offsite Autotuning Approach -- Performance Model Driven Autotuning Applied to Parallel Explicit ODE Methods