2 papers
cs.SE2026
Don't Regenerate, Debug: A Domain-Specific Agent for Repairing Near-Miss Hardware Operators
Yansong Sun, Shenxiu Wu, Siyuan Chen +6
Kernel generation for hardware accelerators such as GPUs and NPUs has become a proving ground for large language models (LLMs), and state-of-the-art systems raise correctness throu…
cs.DC2026
AscendCraft: Automatic Ascend NPU Kernel Generation via DSL-Guided Transcompilation
Zhongzhen Wen, Shudi Shao, Zhong Li +4
The performance of deep learning models critically depends on efficient kernel implementations, yet developing high-performance kernels for specialized accelerators remains time-co…