Showing cs.PLShow all
2 papers · 1 filter
cs.PL2025
Optimal Software Pipelining and Warp Specialization for Tensor Core GPUs
Rupanshu Soi, Rohan Yadav, Fredrik Kjolstad +4
GPU architectures have continued to grow in complexity, with recent incarnations introducing increasingly powerful fixed-function units for matrix multiplication and data movement…
cs.PL2025
Task-Based Tensor Computations on Modern GPUs
Rohan Yadav, Michael Garland, Alex Aiken +1
Domain-specific, fixed-function units are becoming increasingly common in modern processors. As the computational demands of applications evolve, the capabilities and programming i…