1 paper
Weikai Xu, Meng Li, Shuzhang Zhong +7
The Mixture-of-Experts (MoE) models have emerged as the state-of-the-art paradigm for scaling up large language models (LLMs) without proportionally increased computational cost. H…