1 citations · 2 across the 11 of their papers we have counts for
1 paper · 1 filter
Yuhao Qing, Guichao Zhu, Lintian Lei +9
Mixture-of-Experts (MoE) scales large language models cost-effectively, but expert-parallel training suffers severe straggler effects from skewed expert loads. Current systems freq…