Loo.py: transformation-based code generation for GPUs and CPUs
arXiv:1405.7470 · doi:10.1145/2627373.2627387
Abstract
Today's highly heterogeneous computing landscape places a burden on programmers wanting to achieve high performance on a reasonably broad cross-section of machines. To do so, computations need to be expressed in many different but mathematically equivalent ways, with, in the worst case, one variant per target machine. Loo.py, a programming system embedded in Python, meets this challenge by defining a data model for array-style computations and a library of transformations that operate on this model. Offering transformations such as loop tiling, vectorization, storage management, unrolling, instruction-level parallelism, change of data layout, and many more, it provides a convenient way to capture, parametrize, and re-unify the growth among code variants. Optional, deep integration with numpy and PyOpenCL provides a convenient computing environment where the transition from prototype to high-performance implementation can occur in a gradual, machine-assisted form.
References in corpus (3)
Cited by in corpus (16)
- Tensor Comprehensions: Framework-Agnostic High-Performance Machine Learning Abstractions
- TVM: An Automated End-to-End Optimizing Compiler for Deep Learning
- Fast matrix-free evaluation of discontinuous Galerkin finite element operators
- Loo.py: transformation-based code generation for GPUs and CPUs
- Gauge preheating with full general relativity
- A study of vectorization for matrix-free finite element methods
- A shared compilation stack for distributed-memory parallelism in stencil DSLs
- ConfuciuX: Autonomous Hardware Resource Assignment for DNN Accelerators using Reinforcement Learning
- Using SIMD and SIMT vectorization to evaluate sparse chemical kinetic Jacobian matrices and thermochemical source terms
- Loo.py: From Fortran to performance via transformation and substitution rules
- FusionStitching: Deep Fusion and Code Generation for Tensorflow Computations on GPUs
- Array Program Transformation with Loo.py by Example: High-Order Finite Elements
- A mechanism for balancing accuracy and scope in cross-machine black-box GPU performance modeling
- A Unified, Hardware-Fitted, Cross-GPU Performance Model
- Teaching An Old Dog New Tricks: Porting Legacy Code to Heterogeneous Compute Architectures With Automated Code Translation
- Investigating the OPS intermediate representation to target GPUs in the Devito DSL