3 papers
cs.SE2026
RealisticTritonBench: A Benchmark for Triton-Kernel Generation in Real-World AI Frameworks
Jinjun Huang, Zhongzhen Wen, Tongtong Xu +3
In modern AI frameworks, GPU kernels are key to overall system performance. Combining usability, portability, and near-handwritten CUDA performance, Triton is widely adopted for im…
cs.DC2026
AscendCraft: Automatic Ascend NPU Kernel Generation via DSL-Guided Transcompilation
Zhongzhen Wen, Shudi Shao, Zhong Li +4
The performance of deep learning models critically depends on efficient kernel implementations, yet developing high-performance kernels for specialized accelerators remains time-co…
cs.DC2025
MultiKernelBench: A Multi-Platform Benchmark for Kernel Generation
Zhongzhen Wen, Yinghui Zhang, Zhong Li +3
The automatic generation of deep learning (DL) kernels using large language models (LLMs) has emerged as a promising approach to reduce the manual effort and hardware-specific expe…