7 citations · 14 across the 4 of their papers we have counts for
1 paper · 1 filter
Xiaonan Nie, Qibin Liu, Fangcheng Fu +6
Larger transformer models always perform better on various tasks but require more costs to scale up the model size. To efficiently enlarge models, the mixture-of-experts (MoE) arch…