1 paper
Weihao Zhu, Long Shi, Kang Wei +4
As an enabling architecture of Large Models (LMs), Mixture of Experts (MoE) has become prevalent thanks to its sparsely-gated mechanism, which lowers computational overhead while m…