2 citations · 4 across the 5 of their papers we have counts for
1 paper · 1 filter
Chang Chen, Min Li, Zhihua Wu +2
Sparsely gated Mixture-of-Expert (MoE) has demonstrated its effectiveness in scaling up deep neural networks to an extreme scale. Despite that numerous efforts have been made to im…