1 paper
Xiao Shi, Yingying Sun, Jiangsu Du +2
As MoE models scale to hundreds of experts, placement and pruning decisions increasingly dictate communication volume, affecting the performance of distributed inference across GPU…