5 citations · 7 across the 4 of their papers we have counts for
4 papers
On the Duality of Task and Actor Programming Models
Rohan Yadav, Joseph Guman, Sean Treichler +4
Programming models for distributed and heterogeneous machines are rapidly growing in popularity to meet the demands of modern workloads. Task and actor models are common choices th…
Task-Based Tensor Computations on Modern GPUs
Rohan Yadav, Michael Garland, Alex Aiken +1
Domain-specific, fixed-function units are becoming increasingly common in modern processors. As the computational demands of applications evolve, the capabilities and programming i…
Stream-K: Work-centric Parallel Decomposition for Dense Matrix-Matrix Multiplication on the GPU
Muhammad Osama, Duane Merrill, Cris Cecka +2
We introduce Stream-K, a work-centric parallelization of matrix multiplication (GEMM) and related computations in dense linear algebra. Whereas contemporary decompositions are prim…
Efficient Sparsely Activated Transformers
Salar Latifi, Saurav Muralidharan, Michael Garland
Transformer-based neural networks have achieved state-of-the-art task performance in a number of machine learning domains including natural language processing and computer vision.…