4 papers · 1 filter
Stochastic Polyak Stepsize with a Moving Target
Robert M. Gower, Aaron Defazio, Michael Rabbat
We propose a new stochastic gradient method called MOTAPS (Moving Targetted Polyak Stepsize) that uses recorded past loss values to compute adaptive stepsizes. MOTAPS can be seen a…
Variance-Reduced Methods for Machine Learning
Robert M. Gower, Mark Schmidt, Francis Bach +1
Stochastic optimization lies at the heart of machine learning, and its cornerstone is stochastic gradient descent (SGD), a method introduced over 60 years ago. The last 8 years hav…
Almost sure convergence rates for Stochastic Gradient Descent and Stochastic Heavy Ball
Othmane Sebbouh, Robert M. Gower, Aaron Defazio
We study stochastic gradient descent (SGD) and the stochastic heavy ball method (SHB, otherwise known as the momentum method) for the general stochastic approximation problem. For…
SGD: General Analysis and Improved Rates
Robert Mansel Gower, Nicolas Loizou, Xun Qian +3
We propose a general yet simple theorem describing the convergence of SGD under the arbitrary sampling paradigm. Our theorem describes the convergence of an infinite array of varia…