3 papers
cs.LG2026
Does SGD Seek Flatness or Sharpness? An Exactly Solvable Model
Yizhou Xu, Pierfrancesco Beneventano, Isaac Chuang +1
A large body of theory and empirical work hypothesizes a connection between the flatness of a neural network's loss landscape during training and its performance. However, there ha…
cs.LG2025
Gradient Descent Converges Linearly to Flatter Minima than Gradient Flow in Shallow Linear Networks
Pierfrancesco Beneventano, Blake Woodworth
We study the gradient descent (GD) dynamics of a depth-2 linear neural network with a single input and output. We show that GD converges at an explicit linear rate to a global mini…
cs.LG2024
Edge of Stochastic Stability: Revisiting the Edge of Stability for SGD
Arseniy Andreyev, Pierfrancesco Beneventano
Recent findings by Cohen et al., 2021, demonstrate that when training neural networks using full-batch gradient descent with a step size of , the largest eigenvalue o…