Showing cs.AIShow all
2 papers · 1 filter
cs.AI2026
Hawk: Harnessing Hardware-Aware Knowledge for High-Performance NPU Kernel Generation
Junyi Wen, Ruiyan Zhuang, Yongjia Xu +7
Developing high-performance kernels for Neural Processing Units (NPUs) is a critical industry bottleneck, requiring developers to manually navigate implicit hardware constraints an…
cs.AI2025
Fate: Fast Edge Inference of Mixture-of-Experts Models via Cross-Layer Gate
Zhiyuan Fang, Zicong Hong, Yuegui Huang +5
Large Language Models (LLMs) have demonstrated impressive performance across various tasks, and their application in edge scenarios has attracted significant attention. However, sp…