27 citations · 59 across the 8 of their papers we have counts for
1 paper · 1 filter
Lu Zhao, Rong Shi, Shaoqing Zhang +21
The training of large-scale Mixture of Experts (MoE) models faces a critical memory bottleneck due to severe load imbalance caused by dynamic token routing. This imbalance leads to…