133 citations · 651 across the 52 of their papers we have counts for
1 paper · 1 filter
Xin Chen, Hengheng Zhang, Xiaotao Gu +3
The Mixture of Experts (MoE) model becomes an important choice of large language models nowadays because of its scalability with sublinear computational complexity for training and…