1 paper
Hansung Kim, Ruohan Richard Yan, Joshua You +2
Modern GPUs incorporate specialized matrix units such as Tensor Cores to accelerate GEMM operations, which are central to deep learning workloads. However, existing matrix unit des…