6 citations · 6 across the 1 of their papers we have counts for
1 paper
Xun Wu, Shaohan Huang, Wenhui Wang +1
Sparse Mixtures of Experts (SMoE) scales model capacity without significant increases in training and inference costs, but exhibits the following two issues: (1) Low expert activat…