activity
20182022
most citedLocal SGD: Unified Theory and New Efficient Methods

10 citations · 17 across the 3 of their papers we have counts for

collaborators

15 papers

math.OC2022

Distributed Methods with Absolute Compression and Error Compensation

Marina Danilova, Eduard Gorbunov

Distributed optimization methods are often applied to solving huge-scale problems like training neural networks with millions and even billions of parameters. In such applications,…

cs.LG20227 cited

3PC: Three Point Compressors for Communication-Efficient Distributed Training and a Better Theory for Lazy Aggregation

Peter Richtárik, Igor Sokolov, Ilyas Fatkhullin +3

We propose and study a new class of gradient communication mechanisms for communication-efficient training -- three point compressors (3PC) -- as well as efficient distributed nonc…

cs.LG202010 cited

Local SGD: Unified Theory and New Efficient Methods

Eduard Gorbunov, Filip Hanzely, Peter Richtárik

We present a unified framework for analyzing local SGD methods in the convex and strongly convex regimes for distributed/federated training of supervised machine learning models. W…

math.OC2020

Linearly Converging Error Compensated SGD

Eduard Gorbunov, Dmitry Kovalev, Dmitry Makarenko +1

In this paper, we propose a unified analysis of variants of distributed SGD with arbitrary compressions and delayed updates. Our framework is general enough to cover different vari…

math.OC2020

Stochastic Optimization with Heavy-Tailed Noise via Accelerated Gradient Clipping

Eduard Gorbunov, Marina Danilova, Alexander Gasnikov

In this paper, we propose a new accelerated stochastic first-order method called clipped-SSTM for smooth convex stochastic optimization with heavy-tailed distributed noise in stoch…

math.OC2019

Optimal Decentralized Distributed Algorithms for Stochastic Convex Optimization

Eduard Gorbunov, Darina Dvinskikh, Alexander Gasnikov

We consider stochastic convex optimization problems with affine constraints and develop several methods using either primal or dual approach to solve it. In the primal case, we use…