LIKWID: Lightweight Performance Tools
arXiv:1104.4874 · doi:10.1109/ICPPW.2010.38
Abstract
Exploiting the performance of today's microprocessors requires intimate knowledge of the microarchitecture as well as an awareness of the ever-growing complexity in thread and cache topology. LIKWID is a set of command line utilities that addresses four key problems: Probing the thread and cache topology of a shared-memory node, enforcing thread-core affinity on a program, measuring performance counter metrics, and microbenchmarking for reliable upper performance bounds. Moreover, it includes a mpirun wrapper allowing for portable thread-core affinity in MPI and hybrid MPI/threaded applications. To demonstrate the capabilities of the tool set we show the influence of thread affinity on performance using the well-known OpenMP STREAM triad benchmark, use hardware counter tools to study the performance of a stencil code, and finally show how to detect bandwidth problems on ccNUMA-based compute nodes.
12 pages
References in corpus (2)
Cited by in corpus (63)
- A Recursive Algebraic Coloring Technique for Hardware-Efficient Symmetric Sparse Matrix-Vector Multiplication
- Fast matrix-free evaluation of discontinuous Galerkin finite element operators
- Exploring performance and power properties of modern multicore chips via simple machine models
- Quantifying performance bottlenecks of stencil computations using the Execution-Cache-Memory model
- Predictive Performance Modeling for Distributed Computing using Black-Box Monitoring and Machine Learning
- Analytical Characterization and Design Space Exploration for Optimization of CNNs
- High-performance implementation of Chebyshev filter diagonalization for interior eigenvalue computations
- Efficient multicore-aware parallelization strategies for iterative stencil computations
- nanoBench: A Low-Overhead Tool for Running Microbenchmarks on x86 Systems
- From Facility to Application Sensor Data: Modular, Continuous and Holistic Monitoring with DCDB
- Comparing the Performance of Different x86 SIMD Instruction Sets for a Medical Imaging Application on Modern Multi- and Manycore Chips
- Memory Performance of AMD EPYC Rome and Intel Cascade Lake SP Server Processors
- Kerncraft: A Tool for Analytic Performance Modeling of Loop Kernels
- Low-Latency Software Polar Decoders
- Pushing the limits for medical image reconstruction on recent standard multicore processors
- ECM modeling and performance tuning of SpMV and Lattice QCD on A64FX
- Lattice Boltzmann Benchmark Kernels as a Testbed for Performance Analysis
- Performance Modeling of Streaming Kernels and Sparse Matrix-Vector Multiplication on A64FX
- An Efficient ADER-DG Local Time Stepping Scheme for 3D HPC Simulation of Seismic Waves in Poroelastic Media
- Automatic Loop Kernel Analysis and Performance Modeling With Kerncraft
- Performance Engineering of the Kernel Polynomial Method on Large-Scale CPU-GPU Systems
- Efficient implementation of modern entropy stable and kinetic energy preserving discontinuous Galerkin methods for conservation laws
- A Mess of Memory System Benchmarking, Simulation and Application Profiling
- Best practices for HPM-assisted performance engineering on modern multicore processors
- Bandwidth-Aware Page Placement in NUMA
- Using analog computers in today's largest computational challenges
- Studies on the energy and deep memory behaviour of a cache-oblivious, task-based hyperbolic PDE solver
- Fast Tensor Product Schwarz Smoothers for High-Order Discontinuous Galerkin Methods
- Chebyshev Filter Diagonalization on Modern Manycore Processors and GPGPUs
- Performance Characterization of Multi-threaded Graph Processing Applications on Intel Many-Integrated-Core Architecture
- kEDM: A Performance-portable Implementation of Empirical Dynamic Modeling using Kokkos
- Complex additive geometric multilevel solvers for Helmholtz equations on spacetrees
- Improved vectorization of OpenCV algorithms for RISC-V CPUs
- Stop talking to me -- a communication-avoiding ADER-DG realisation
- Performance of the low-rank tensor-train SVD (TT-SVD) for large dense tensors on modern multi-core CPUs
- Accurate Measurement of Application-level Energy Consumption for Energy-Aware Large-Scale Simulations
- Supercomputing with MPI meets the Common Workflow Language standards: an experience report
- Performance Engineering for a Medical Imaging Application on the Intel Xeon Phi Accelerator
- Analytic Performance Modeling and Analysis of Detailed Neuron Simulations
- Observing the Invisible: Live Cache Inspection for High-Performance Embedded Systems
- Validation of hardware events for successful performance pattern identification in High Performance Computing
- Mixed-mode implementation of PETSc for scalable linear algebra on multi-core processors
- Energy of Computing on Multicore CPUs: Predictive Models and Energy Conservation Law
- ThirstyFLOPS: Water Footprint Modeling and Analysis Toward Sustainable HPC Systems
- SPEChpc 2021 Benchmarks on Ice Lake and Sapphire Rapids Infiniband Clusters: A Performance and Energy Case Study
- Modernizing an Operational Real-time Tsunami Simulator to Support Diverse Hardware Platforms
- TEEMon: A continuous performance monitoring framework for TEEs
- Integrating State of the Art Compute, Communication, and Autotuning Strategies to Multiply the Performance of the Application Programm CPMD for Ab Initio Molecular Dynamics Simulations
- Calculating Software's Energy Use and Carbon Emissions: A Survey of the State of Art, Challenges, and the Way Ahead
- Toward Efficient In-memory Data Analytics on NUMA Systems
- Vectorization of Gradient Boosting of Decision Trees Prediction in the CatBoost Library for RISC-V Processors
- An analytic performance model for overlapping execution of memory-bound loop kernels on multicore CPUs
- MT4G: A Tool for Reliable Auto-Discovery of NVIDIA and AMD GPU Compute and Memory Topologies
- Hardware-Oriented Krylov Methods for High-Performance Computing
- MCompiler: A Synergistic Compilation Framework
- A Hermite-like basis for faster matrix-free evaluation of interior penalty discontinuous Galerkin operators
- A multiresolution Discrete Element Method for triangulated objects with implicit timestepping
- Optimizing Large-Scale ODE Simulations
- Cache Blocking of Distributed-Memory Parallel Matrix Power Kernels
- Marrying Many-core Accelerators and InfiniBand for a New Commodity Processor
- Modern Multicore CPUs are not Energy Proportional: Opportunity for Bi-objective Optimization for Performance and Energy
- Double-precision FPUs in High-Performance Computing: an Embarrassment of Riches?
- Using performance analysis tools for parallel-in-time integrators -- Does my time-parallel code do what I think it does?