paper

Optimal Two-Step Stepsize Schedule for Stochastic Gradient Methods

arXiv:2608.15035

Abstract

Structured nonconstant large stepsizes can improve the convergence of gradient descent in the deterministic setting. However, in stochastic optimization, aggressive stepsizes can amplify oracle noise and hinder the convergence of stochastic gradient methods. We characterize the globally optimal two-step stepsize schedule for stochastic gradient methods applied to strongly convex and smooth functions, assuming access only to unbiased stochastic gradient estimates with finite support and bounded variance. The optimal schedule depends on the ratio of the initial optimality gap to the noise level and exhibits several distinct regimes. As the influence of stochastic noise diminishes, the optimal two-step stepsizes become larger, reflecting a balance between the benefits of faster iterate convergence and the perturbations induced by stochastic noise.

23 pages, 2 figures