1.7k citations · 1.9k across the 15 of their papers we have counts for
3 papers · 1 filter
Adaptive Learning Rates for Faster Stochastic Gradient Methods
Samuel Horváth, Konstantin Mishchenko, Peter Richtárik
In this work, we propose new adaptive step size strategies that improve several stochastic gradient methods. Our first method (StoPS) is based on the classical Polyak step size (Po…
Variance Reduced ProxSkip: Algorithm, Theory and Application to Federated Learning
Grigory Malinovsky, Kai Yi, Peter Richtárik
We study distributed optimization methods based on the {\em local training (LT)} paradigm: achieving communication efficiency by performing richer local gradient-based training on…
Communication Acceleration of Local Gradient Methods via an Accelerated Primal-Dual Algorithm with Inexact Prox
Abdurakhmon Sadiev, Dmitry Kovalev, Peter Richtárik
Inspired by a recent breakthrough of Mishchenko et al (2022), who for the first time showed that local gradient steps can lead to provable communication acceleration, we propose an…