60 citations · 60 across the 1 of their papers we have counts for
1 paper
Yanqi Zhou, Tao Lei, Hanxiao Liu +7
Sparsely-activated Mixture-of-experts (MoE) models allow the number of parameters to greatly increase while keeping the amount of computation for a given token or a given sample un…