3 papers
cs.DC2026
TerraceMoE: A Cost Model for Hierarchical MoE All-to-All Communication
Weicheng Xue, Bingqiang Wang, Li Yuan +2
Hierarchical two-hop dispatch can reduce slow-fabric traffic in expert-parallel Mixture-of-Experts training, but it adds a second collective and an arrival-side operator chain. We…
cs.DC2026
Ascend to Science: Exploration of AI Chips for Scientific Computing
Weicheng Xue, Kai Yang, Yongxiang Liu +5
The rapid rise of AI-oriented accelerators has reshaped compute systems around low-precision tensor engines, raising a practical question for the HPC community: under what conditio…
cs.AI2026
AscendKernelGen: A Systematic Study of LLM-Based Kernel Generation for Neural Processing Units
Xinzi Cao, Jianyang Zhai, Pengfei Li +17
To meet the ever-increasing demand for computational efficiency, Neural Processing Units (NPUs) have become critical in modern AI infrastructure. However, unlocking their full pote…