32 citations · 32 across the 1 of their papers we have counts for
1 paper
Changho Hwang, Wei Cui, Yifan Xiong +12
Sparsely-gated mixture-of-experts (MoE) has been widely adopted to scale deep learning models to trillion-plus parameters with fixed computational cost. The algorithmic performance…