6 papers
Dynamics of Gradient Descent with Large Step Size Near a Manifold of Flat Minima
Lachlan Ewen MacDonald, René Vidal
An important quantity in the theory of gradient descent (GD) is the \emph{sharpness}, defined as the largest eigenvalue of the objective Hessian. Classical analyses typically requi…
SGD at the Edge of Stability: Stochastic Stabilization with Large Learning Rates
Konstantinos Emmanouilidis, Lachlan MacDonald, Salma Tarmoun +1
Modern deep learning has been shown to operate at the edge of stability, routinely using learning rates far larger than those justified by classical optimization theory. Most prior…
Centre manifold theorem for maps along manifolds of fixed points
Lachlan Ewen MacDonald
We prove a centre manifold theorem for a map along a manifold-with-boundary of fixed points, and provide an application to the study of gradient descent with large step size on two…
Convergence Rates for Gradient Descent on the Edge of Stability in Overparametrised Least Squares
Lachlan Ewen MacDonald, Hancheng Min, Leandro Palma +3
Classical optimisation theory guarantees monotonic objective decrease for gradient descent (GD) when employed in a small step size, or ``stable", regime. In contrast, gradient desc…
Disintegration theorem for multifunctions, with applications to empirical Wasserstein distances and average-case statistical bounds
James Allen Fill, Lachlan Ewen MacDonald
We prove a generalisation of the disintegration theorem to the setting of multifunctions between Polish probability spaces. Whereas the classical disintegration theorem guarantees…
Understanding the Learning Dynamics of LoRA: A Gradient Flow Perspective on Low-Rank Adaptation in Matrix Factorization
Ziqing Xu, Hancheng Min, Lachlan Ewen MacDonald +4
Despite the empirical success of Low-Rank Adaptation (LoRA) in fine-tuning pre-trained models, there is little theoretical understanding of how first-order methods with carefully c…