2 citations · 2 across the 1 of their papers we have counts for
1 paper
Chang Chen, Min Li, Zhihua Wu +2
Sparsely gated Mixture-of-Expert (MoE) has demonstrated its effectiveness in scaling up deep neural networks to an extreme scale. Despite that numerous efforts have been made to im…