2 papers
cs.DC2025
Breaking (Global) Barriers in Parallel Stochastic Optimization with Wait-Avoiding Group Averaging
Shigang Li, Tal Ben-Nun, Giorgi Nadiradze +4
Deep learning at scale is dominated by communication time. Distributing samples across nodes usually yields the best performance, but poses scaling challenges due to global informa…
cs.DC2025
Taming Unbalanced Training Workloads in Deep Learning with Partial Collective Operations
Shigang Li, Tal Ben-Nun, Salvatore Di Girolamo +2
Load imbalance pervasively exists in distributed deep learning training systems, either caused by the inherent imbalance in learned tasks or by the system itself. Traditional synch…