1 paper
Juntong Wu, Yifei Liu, Junyi Chen +6
Mixture-of-Experts (MoE) models scale capacity for strong quality while keeping per-token compute bounded through sparse expert activation. Yet low-latency MoE serving is increasin…