32 citations · 32 across the 1 of their papers we have counts for
1 paper
Stanislaw Jastrzebski, Maciej Szymczak, Stanislav Fort +4
The early phase of training of deep neural networks is critical for their final performance. In this work, we study how the hyperparameters of stochastic gradient descent (SGD) use…