1 paper
Weiwei Chen, Shuang Chen, Lele Li +5
Mixture-of-Experts (MoE) models have become the de facto standard for scaling large language models. To maintain computational efficiency, modern MoE serving systems typically empl…