collaborators

5 papers

cs.DC2026

HyperParallel-MoE: Multi-Core Interleaved Scheduling for Fast MoE Training on Ascend NPUs

Zewen Jin, Congkun Ai, Guangpeng Zhang +7

Modern Mixture-of-Experts (MoE) models increasingly rely on large-scale AI accelerator clusters for efficient training. Ascend NPUs expose heterogeneous on-chip compute resources,…

cs.PL2026

DVM: A Bytecode Virtual Machine Approach for Dynamic Tensor Computation

Jingzhi Fang, Xiong Gao, Renwei Zhang +6

Dynamism is common in AI computation, e.g., the dynamic tensor shapes and the dynamic control flows in models. Due to the long compilation time, existing runtime compilation damage…

cs.DC2026

HyperParallel: A Supernode-Affinity AI Framework

Xin Zhang, Beilei Sun, Teng Su +4

The emergence of large-scale, sparse, multimodal, and agentic AI models has coincided with a shift in hardware toward supernode architectures that integrate hundreds to thousands o…

cs.DC2026

HyperOffload: Graph-Driven Hierarchical Memory Management for Large Language Models on SuperNode Architectures

Fangxin Liu, Qinghua Zhang, Hanjing Shen +5

The rapid evolution of Large Language Models (LLMs) towards long-context reasoning and sparse architectures has pushed memory requirements far beyond the capacity of individual dev…

cs.AI2025

AKG kernel Agent: A Multi-Agent Framework for Cross-Platform Kernel Synthesis

Jinye Du, Quan Yuan, Zuyao Zhang +16

Modern AI models demand high-performance computation kernels. The growing complexity of LLMs, multimodal architectures, and recommendation systems, combined with techniques like sp…