1 paper
Philip M. Long, Peter L. Bartlett
Recent experiments have shown that, often, when training a neural network with gradient descent (GD) with a step size I^⋅, the operator norm of the Hessian of the loss grows until…