1 paper
Chao Jin, Xinming Wei, Yinmin Zhong +6
Load imbalance is a long-standing challenge in Mixture-of-Experts (MoE) training and is exacerbated in reinforcement learning (RL) for LLMs, where hot experts can shift frequently…