16 citations · 82 across the 29 of their papers we have counts for
Showing 2020 · math.OCShow all
2 papers · 2 filters
math.OC2020
SGD for Structured Nonconvex Functions: Learning Rates, Minibatching and Interpolation
Robert M. Gower, Othmane Sebbouh, Nicolas Loizou
Stochastic Gradient Descent (SGD) is being used routinely for optimizing non-convex functions. Yet, the standard convergence theory for SGD in the smooth non-convex setting gives a…
math.OC2020
Stochastic Polyak Step-size for SGD: An Adaptive Learning Rate for Fast Convergence
Nicolas Loizou, Sharan Vaswani, Issam Laradji +1
We propose a stochastic variant of the classical Polyak step-size (Polyak, 1987) commonly used in the subgradient method. Although computing the Polyak step-size requires knowledge…