Search papers, labs, and topics across Lattice.
This paper identifies a critical bottleneck in Mixture of Experts (MoE) architectures caused by traditional round-robin (RR) scheduling, which leads to an exponential incast phenomenon in MoE traffic. To address this issue, the authors propose a proactive fair scheduling framework specifically designed for MoE workloads, which prevents fabric oversubscription and enhances performance. Extensive simulations reveal that this new framework not only eliminates incast but also achieves near-100% link utilization and significantly reduces Collective Completion Time (CCT).
Traditional round-robin scheduling in MoE architectures can lead to an exponential incast problem, but a new proactive scheduling framework effectively eliminates this bottleneck.
Mixture of Experts (MoE) architectures have become key to large language models; however, their typical round-robin (RR) scheduling introduces significant bottlenecks. In this paper, we demonstrate that RR causes a previously-undiscovered exponential incast phenomenon with MoE traffic. We propose an alternative proactive fair scheduling framework tailored for MoE workloads, which effectively prevents fabric oversubscription. We also outline how it can be implemented in NICs. Finally, through extensive simulations with real and synthetic workloads, we demonstrate that this framework consistently eliminates incast, maintains a near-100% link utilization, and reduces Collective Completion Time (CCT).