2 citations · 2 across the 1 of their papers we have counts for
1 paper · 1 filter
Mengting Ai, Tianxin Wei, Yifan Chen +7
Mixture-of-Experts (MoE) Transformer, the backbone architecture of multiple phenomenal language models, leverages sparsity by activating only a fraction of model parameters for eac…