2 citations · 3 across the 3 of their papers we have counts for
1 paper · 1 filter
Yongji Wu, Wenjie Qu, Xueshen Liu +10
Sparsely-activated Mixture-of-Experts (MoE) architecture has increasingly been adopted to further scale large language models (LLMs). However, frequent failures still pose signific…