Efficient multicore-aware parallelization strategies for iterative stencil computations
arXiv:1004.1741 · doi:10.1016/j.jocs.2011.01.010
Abstract
Stencil computations consume a major part of runtime in many scientific simulation codes. As prototypes for this class of algorithms we consider the iterative Jacobi and Gauss-Seidel smoothers and aim at highly efficient parallel implementations for cache-based multicore architectures. Temporal cache blocking is a known advanced optimization technique, which can reduce the pressure on the memory bus significantly. We apply and refine this optimization for a recently presented temporal blocking strategy designed to explicitly utilize multicore characteristics. Especially for the case of Gauss-Seidel smoothers we show that simultaneous multi-threading (SMT) can yield substantial performance improvements for our optimized algorithm.
15 pages, 10 figures
References in corpus (3)
Cited by in corpus (5)
- LIKWID: Lightweight Performance Tools
- Real-space density functional theory on graphical processing units: computational approach and comparison to Gaussian basis set methods
- LIKWID: A lightweight performance-oriented tool suite for x86 multicore environments
- Optimizing ccNUMA locality for task-parallel execution under OpenMP and TBB on multicore-based systems
- Tight Bounds for Low Dimensional Star Stencils in the Parallel External Memory Model