2 papers
cs.LG2026
DynamiQ: Accelerating Gradient Synchronization using Compressed Multi-hop All-reduce
Wenchen Han, Shay Vargaftik, Michael Mitzenmacher +1
Multi-hop all-reduce is the de facto backbone of large model training. As the training scale increases, the network often becomes a bottleneck, motivating the reduction of the volu…
cs.DC2025
Bounded Memory in Distributed Networks
Ran Ben Basat, Keren Censor-Hillel, Yi-Jun Chang +3
The recent advent of programmable switches makes distributed algorithms readily deployable in real-world datacenter networks. However, there are still gaps between theory and pract…