Iterate averaging as regularization for stochastic gradient descent
arXiv:1802.08009
Abstract
We propose and analyze a variant of the classic Polyak-Ruppert averaging scheme, broadly used in stochastic gradient methods. Rather than a uniform average of the iterates, we consider a weighted average, with weights decaying in a geometric fashion. In the context of linear least squares regression, we show that this averaging scheme has a the same regularizing effect, and indeed is asymptotically equivalent, to ridge regression. In particular, we derive finite-sample bounds for the proposed approach that match the best known results for regularized stochastic gradient methods.
Cited by in corpus (7)
- The Step Decay Schedule: A Near Optimal, Geometrically Decaying Learning Rate Procedure For Least Squares
- The Implicit Regularization of Stochastic Gradient Flow for Least Squares
- Beating SGD Saturation with Tail-Averaging and Minibatching
- Server Averaging for Federated Learning
- Efficient and Robust Algorithms for Adversarial Linear Contextual Bandits
- On the connections between algorithmic regularization and penalization for convex losses
- Distribution-Dependent Analysis of Gibbs-ERM Principle