1 paper
Rui Kong, Yuanchun Li, Qingtian Feng +5
Mixture of experts (MoE) is a popular technique to improve capacity of Large Language Models (LLMs) with conditionally-activated parallel experts. However, serving MoE models on me…