1 paper
Arseniy Andreyev, Pierfrancesco Beneventano
Recent findings by Cohen et al., 2021, demonstrate that when training neural networks using full-batch gradient descent with a step size of η, the largest eigenvalue λmax o…