52 citations · 57 across the 3 of their papers we have counts for
4 papers
CuTe Layout Representation and Algebra
Cris Cecka
Modern architectures for high-performance computing and deep learning increasingly incorporate specialized tensor instructions, including tensor cores for matrix multiplication and…
Stream-K: Work-centric Parallel Decomposition for Dense Matrix-Matrix Multiplication on the GPU
Muhammad Osama, Duane Merrill, Cris Cecka +2
We introduce Stream-K, a work-centric parallelization of matrix multiplication (GEMM) and related computations in dense linear algebra. Whereas contemporary decompositions are prim…
Tensor Contractions with Extended BLAS Kernels on CPU and GPU
Yang Shi, U. N. Niranjan, Animashree Anandkumar +1
Tensor contractions constitute a key computational ingredient of numerical multi-linear algebra. However, as the order and dimension of tensors grow, the time and space complexitie…
Fourier Based Fast Multipole Method for the Helmholtz Equation
Cris Cecka, Eric Darve
The fast multipole method (FMM) has had great success in reducing the computational complexity of solving the boundary integral form of the Helmholtz equation. We present a formula…