4 papers · 1 filter
Learning Under Delayed Feedback: Implicitly Adapting to Gradient Delays
Rotem Zamir Aviv, Ido Hakimi, Assaf Schuster +1
We consider stochastic convex optimization problems, where several machines act asynchronously in parallel while sharing a common memory. We propose a robust training method for th…
It's Not What Machines Can Learn, It's What We Cannot Teach
Gal Yehuda, Moshe Gabel, Assaf Schuster
Can deep neural networks learn to solve any task, and in particular problems of high complexity? This question attracts a lot of interest, with recent works tackling computationall…
Gap Aware Mitigation of Gradient Staleness
Saar Barkai, Ido Hakimi, Assaf Schuster
Cloud computing is becoming increasingly popular as a platform for distributed training of deep neural networks. Synchronous stochastic gradient descent (SSGD) suffers from substan…
Taming Momentum in a Distributed Asynchronous Environment
Ido Hakimi, Saar Barkai, Moshe Gabel +1
Although distributed computing can significantly reduce the training time of deep neural networks, scaling the training process while maintaining high efficiency and final accuracy…