1 paper
Nehal Baganal Krishna, Anam Tahir, Firas Khamis +3
Large-scale training for distributed Machine Learning can cause congestion at bottleneck switch ports, leading to model staleness through update losses. This is particularly detrim…