1 paper
En-Ming Huang, An-Cheng Chang, Bai-Cheng Jeng +2
Mixture-of-Experts (MoE) inference consists of sparse expert GEMMs whose shapes vary with the runtime routing distribution. Existing serving systems typically select fused-MoE kern…