uops.info: Characterizing Latency, Throughput, and Port Usage of Instructions on Intel Microarchitectures
arXiv:1810.04610 · doi:10.1145/3297858.3304062
Abstract
Modern microarchitectures are some of the world's most complex man-made systems. As a consequence, it is increasingly difficult to predict, explain, let alone optimize the performance of software running on such microarchitectures. As a basis for performance predictions and optimizations, we would need faithful models of their behavior, which are, unfortunately, seldom available. In this paper, we present the design and implementation of a tool to construct faithful models of the latency, throughput, and port usage of x86 instructions. To this end, we first discuss common notions of instruction throughput and port usage, and introduce a more precise definition of latency that, in contrast to previous definitions, considers dependencies between different pairs of input and output operands. We then develop novel algorithms to infer the latency, throughput, and port usage based on automatically-generated microbenchmarks that are more accurate and precise than existing work. To facilitate the rapid construction of optimizing compilers and tools for performance prediction, the output of our tool is provided in a machine-readable format. We provide experimental results for processors of all generations of Intel's Core architecture, i.e., from Nehalem to Coffee Lake, and discuss various cases where the output of our tool differs considerably from prior work.
Cited by in corpus (11)
- nanoBench: A Low-Overhead Tool for Running Microbenchmarks on x86 Systems
- ALTO: Adaptive Linearized Storage of Sparse Tensors
- SLEEF: A Portable Vectorized Library of C Standard Mathematical Functions
- uiCA: Accurate Throughput Prediction of Basic Blocks on Recent Intel Microarchitectures
- Automatic Throughput and Critical Path Analysis of x86 and ARM Assembly Kernels
- SpectreRewind: Leaking Secrets to Past Instructions
- Facile: Fast, Accurate, and Interpretable Basic-Block Throughput Prediction
- Transcoding Unicode Characters with AVX-512 Instructions
- Programming with Neural Surrogates of Programs
- Batched Ranged Random Integer Generation
- CACHE SNIPER : Accurate timing control of cache evictions