High-Performance Tensor Contraction without Transposition
arXiv:1607.00291 · doi:10.1137/16M108968X
Abstract
Tensor computations--in particular tensor contraction (TC)--are important kernels in many scientific computing applications. Due to the fundamental similarity of TC to matrix multiplication (MM) and to the availability of optimized implementations such as the BLAS, tensor operations have traditionally been implemented in terms of BLAS operations, incurring both a performance and a storage overhead. Instead, we implement TC using the flexible BLIS framework, which allows for transposition (reshaping) of the tensor to be fused with internal partitioning and packing operations, requiring no explicit transposition operations or additional workspace. This implementation, TBLIS, achieves performance approaching that of MM, and in some cases considerably higher than that of traditional TC. Our implementation supports multithreading using an approach identical to that used for MM in BLIS, with similar performance characteristics. The complexity of managing tensor-to-matrix transformations is also handled automatically in our approach, greatly simplifying its use in scientific applications.
24 pages, 8 figures, uses pgfplots
References in corpus (1)
Cited by in corpus (18)
- Format Abstraction for Sparse Tensor Algebra Compilers
- Tensor renormalization group study of the 3d model
- When gold is not enough: platinum standard of quantum chemistry with cost
- Towards High Performance Relativistic Electronic Structure Modelling: The EXP-T Program Package
- Yet Another Tensor Toolbox for discontinuous Galerkin methods and other applications
- Analytic Gradients of Approximate Coupled Cluster Methods with Quadruple Excitations
- Theory and Implementation of a Novel Stochastic Approach to Coupled Cluster
- Implementing Strassen's Algorithm with CUTLASS on NVIDIA Volta GPUs
- Equation Generator for Equation-of-Motion Coupled Cluster Assisted by Computer Algebra System
- ByteQC: GPU-Accelerated Quantum Chemistry Package for Large-Scale Systems
- Generating coupled cluster code for modern distributed memory tensor software
- Shifting sands of hardware and software in exascale quantum mechanical simulations
- The landscape of software for tensor computations
- a-Tucker: Input-Adaptive and Matricization-Free Tucker Decomposition for Dense Tensors on CPUs and GPUs
- Fast Kronecker Matrix-Matrix Multiplication on GPUs
- GuiTeNet: A graphical user interface for tensor networks
- Supporting mixed-datatype matrix multiplication within the BLIS framework
- A new open-shell CCSDTQ implementation and its application to the basis set convergence of post-CCSDT(Q) corrections in computational thermochemistry