computer networking

Incast-Free MoE Rate-Based Scheduling

arXiv:2607.26340

summary

The paper shows that round-robin scheduling in Mixture of Experts (MoE) models creates an exponential incast problem, and introduces a proactive fair rate‑based scheduling framework that avoids fabric oversubscription, can be realized in NICs, and improves link utilization while reducing collective completion time.

Abstract

Mixture of Experts (MoE) architectures have become key to large language models; however, their typical round-robin (RR) scheduling introduces significant bottlenecks. In this paper, we demonstrate that RR causes a previously-undiscovered exponential incast phenomenon with MoE traffic. We propose an alternative proactive fair scheduling framework tailored for MoE workloads, which effectively prevents fabric oversubscription. We also outline how it can be implemented in NICs. Finally, through extensive simulations with real and synthetic workloads, we demonstrate that this framework consistently eliminates incast, maintains a near-100% link utilization, and reduces Collective Completion Time (CCT).

Topics & keywords

#mixture of experts#incast#scheduling#network fabric#nic implementationround-robin schedulingrate-based schedulingfabric oversubscriptioncollective completion timesimulation