1 paper
Jason Yoo, Rajarshi Saha, Shaowei Zhu +3
Despite rapid progress in LLM-based code generation, writing correct and performant kernels for hardware accelerators remains a key bottleneck in scaling modern ML workloads. We pr…