1 paper
Nikolay Bogoychev, Marcin Junczys-Dowmunt, Kenneth Heafield +1
In order to extract the best possible performance from asynchronous stochastic gradient descent one must increase the mini-batch size and scale the learning rate accordingly. In or…