2 papers
cs.AR2026
Decoding the Skew: Distribution-Aware MoE Inference with Adaptive Kernel Dispatch
En-Ming Huang, An-Cheng Chang, Bai-Cheng Jeng +2
Mixture-of-Experts (MoE) inference consists of sparse expert GEMMs whose shapes vary with the runtime routing distribution. Existing serving systems typically select fused-MoE kern…
cs.DC2025
Efficient CPU-GPU Collaborative Inference for MoE-based LLMs on Memory-Limited Systems
En-Ming Huang, Li-Shang Lin, Chun-Yi Lee
Large Language Models (LLMs) have achieved impressive results across various tasks, yet their high computational demands pose deployment challenges, especially on consumer-grade ha…