1 paper
Seokjin Go, Marko Scrbak, Ephrem Wu +2
In distributed Mixture-of-Experts (MoE) inference, input-dependent token routing interacts with GPU performance variability to create persistent stragglers under synchronized execu…