6 papers
PoLoRA: A Preconditioned Orthogonalized LoRA Optimizer
Nikhil Ghosh, Tetiana Parshakova, Robert M. Gower
Low-rank adaptation (LoRA) makes finetuning large language models cheaper by adding to each weight matrix a trainable low-rank update parameterized as the product of two matrices.…
Non-Euclidean Gradient Descent Operates at the Edge of Stability
Rustem Islamov, Michael Crawshaw, Jeremy Cohen +1
The Edge of Stability (EoS) is a phenomenon where the sharpness (largest eigenvalue) of the Hessian approaches and then hovers near the stability threshold during gradient d…
Step-Size Stability in Stochastic Optimization: A Theoretical Perspective
Fabian Schaipp, Robert M. Gower, Adrien Taylor
We present a theoretical analysis of stochastic optimization methods in terms of their sensitivity with respect to the step size. We identify a key quantity that, for each method,…
Muon Does Not Converge on Convex Lipschitz Functions
Tetiana Parshakova, Ahmed Khaled, Michael Crawshaw +2
Muon and its variants have shown strong empirical performance in a variety of deep learning tasks. Existing convergence analyses of Muon rely on smoothness assumptions, though argu…
Tracking the Median of Gradients with a Stochastic Proximal Point Method
Fabian Schaipp, Guillaume Garrigos, Umut Simsekli +1
There are several applications of stochastic optimization where one can benefit from a robust estimate of the gradient. For example, domains such as distributed learning with corru…
Analysis of an Idealized Stochastic Polyak Method and its Application to Black-Box Model Distillation
Robert M. Gower, Guillaume Garrigos, Nicolas Loizou +3
We provide a general convergence theorem of an idealized stochastic Polyak step size called SPS. Besides convexity, we only assume a local expected gradient bound, that include…