2 papers
cs.DC2026
Beyond Uniform Experts: Cost-Aware Expert Execution for Efficient Multi-Device MoE Inference
Hui Zang, Pengfei Xia, Hong Liu +5
Mixture-of-Experts (MoE) architectures enable language models to achieve unprecedented scale via sparse activation. However, their inference performance is often limited by data mo…
cs.DC2026
FlexNPU: Transparent NPU Virtualization for Dynamic LLM Prefill-Decode Co-location
Jiongjiong Gu, Jianfeng Wang, Zidong Han +19
Modern AI serving increasingly relies on NPUs for conventional inference and large language model serving. However, current NPU deployments commonly expose physical devices directl…