1 paper
Heng Xu, Zhiwei Yu, Chengze Du +5
Training Mixture-of-Experts (MoE) models introduces sparse and highly imbalanced all-to-all communication that dominates iteration time. Conventional load-balancing methods fail to…