8 citations · 12 across the 6 of their papers we have counts for
1 paper · 2 filters
Sajal Dash, Feiyi Wang
Frontier models increasingly adopt Mixture-of-Experts (MoE) architectures to achieve large-model performance at reduced cost. However, training MoE models on HPC platforms is hinde…