1 paper
Xin Chen, Hengheng Zhang, Xiaotao Gu +3
The Mixture of Experts (MoE) model becomes an important choice of large language models nowadays because of its scalability with sublinear computational complexity for training and…