1 paper · 1 filter
Juntong Wu, Jialiang Cheng, Fuyu Lv +2
Mixture-of-Experts (MoE) architectures employ sparse activation to deliver faster training and inference with higher accuracy than dense LLMs. However, in production serving, MoE m…