activity
20182023
most citedStochastic Distributed Learning with Gradient Quantization and Variance Reduction

81 citations · 147 across the 20 of their papers we have counts for

collaborators
Showing 2019 · math.OCShow all

5 papers · 2 filters

math.OC2019

Adaptive Gradient Descent without Descent

Yura Malitsky, Konstantin Mishchenko

We present a strikingly simple proof that two rules are sufficient to automate gradient descent: 1) don't increase the stepsize too fast and 2) don't overstep the local curvature.…

math.OC2019★ 6 cited

MISO is Making a Comeback With Better Proofs and Rates

Xun Qian, Alibek Sailanbayev, Konstantin Mishchenko +1

MISO, also known as Finito, was one of the first stochastic variance reduced methods discovered, yet its popularity is fairly low. Its initial analysis was significantly limited by…

math.OC2019

A Stochastic Decoupling Method for Minimizing the Sum of Smooth and Non-Smooth Functions

Konstantin Mishchenko, Peter Richtárik

We consider the problem of minimizing the sum of three convex functions: i) a smooth function in the form of an expectation or a finite average, ii) a non-smooth function i…

math.OC2019

Revisiting Stochastic Extragradient

Konstantin Mishchenko, Dmitry Kovalev, Egor Shulgin +2

We fix a fundamental issue in the stochastic extragradient method by providing a new sampling strategy that is motivated by approximating implicit updates. Since the existing stoch…

math.OC2019★ 81 cited

Stochastic Distributed Learning with Gradient Quantization and Variance Reduction

Samuel Horváth, Dmitry Kovalev, Konstantin Mishchenko +2

We consider distributed optimization where the objective function is spread among different devices, each sending incremental model updates to a central server. To alleviate the co…