2 citations · 3 across the 3 of their papers we have counts for
1 paper · 1 filter
Sandesh Kamath, Amit Deshpande, K V Subrahmanyam
Learning rate, batch size and momentum are three important hyperparameters in the SGD algorithm. It is known from the work of Jastrzebski et al. arXiv:1711.04623 that large batch s…