2 papers
cs.LG2025
Training Instabilities Induce Flatness Bias in Gradient Descent
Lawrence Wang, Stephen J. Roberts
Classical analyses of gradient descent (GD) define a stability threshold based on the largest eigenvalue of the loss Hessian, often termed sharpness. When the learning rate lies be…
cs.LG2024
Can Stability be Detrimental? Better Generalization through Gradient Descent Instabilities
Lawrence Wang, Stephen J. Roberts
Traditional analyses of gradient descent optimization show that, when the largest eigenvalue of the loss Hessian - often referred to as the sharpness - is below a critical learning…