2 papers
cs.LG2019
Gap Aware Mitigation of Gradient Staleness
Saar Barkai, Ido Hakimi, Assaf Schuster
Cloud computing is becoming increasingly popular as a platform for distributed training of deep neural networks. Synchronous stochastic gradient descent (SSGD) suffers from substan…
cs.LG2019
Taming Momentum in a Distributed Asynchronous Environment
Ido Hakimi, Saar Barkai, Moshe Gabel +1
Although distributed computing can significantly reduce the training time of deep neural networks, scaling the training process while maintaining high efficiency and final accuracy…