From the 1 of 4 linked papers with an AI index.
4 papers
Incast-Free MoE Rate-Based Scheduling
Evyatar Cohen, Jose Yallouz, Alexander Shpiner +3
The paper shows that round-robin scheduling in Mixture of Experts (MoE) models creates an exponential incast problem, and introduces a proactive fair rate‑based scheduling framewor…
Learning-Augmented Scalable Linear Assignment Problem Optimization via Neural Dual Warm-Starts
Ilay Yavlovich, Jad Agbaria, Muhamed Mhamed +2
The Linear Assignment Problem is a fundamental combinatorial optimization task where classical exact solvers ensure optimality but suffer from an bottleneck, w…
Scaling Routers with In-Package Optics and High-Bandwidth Memories
Isaac Keslassy, Ilay Yavlovich, Jose Yallouz +3
This paper aims to apply two major scaling transformations from the computing packaging industry to internet routers: the heterogeneous integration of high-bandwidth memories (HBMs…
Routing for Large ML Models
Ofir Cohen, Jose Yallouz Michael Schapira, Shahar Belkar +1
Training large language models (LLMs), and other large machine learning models, involves repeated communication of large volumes of data across a data center network. The communica…