9 citations · 13 across the 15 of their papers we have counts for
1 paper · 1 filter
Simeng Sun, Roger Waleffe
When training Mixture-of-Experts (MoE) language models with expert parallelism, all-to-all token dispatch and combine collectives can consume a substantial fraction of end-to-end t…