29 citations · 30 across the 8 of their papers we have counts for
1 paper · 2 filters
Yitong Ding, Jiawei Huang, Renyang Guan +7
Modern LLM (large language model) workloads increasingly rely on optimized GPU kernels through hardware-software co-design. These kernels achieve high-performance through fine-grai…