1 paper
Tehila Dahan, Roie Reshef, Sharon Goldstein +1
Asynchronous stochastic gradient descent (SGD) enables scalable distributed training but suffers from gradient staleness. Existing mitigation strategies, such as delay-adaptive lea…