81 citations · 116 across the 8 of their papers we have counts for
9 papers · 1 filter
On Seven Fundamental Optimization Challenges in Machine Learning
Konstantin Mishchenko
Many recent successes of machine learning went hand in hand with advances in optimization. The exchange of ideas between these fields has worked both ways, with machine learning bu…
Random Reshuffling: Simple Analysis with Vast Improvements
Konstantin Mishchenko, Ahmed Khaled, Peter Richtárik
Random Reshuffling (RR) is an algorithm for minimizing finite-sum functions that utilizes iterative gradient descent steps in conjunction with data reshuffling. Often contrasted wi…
Adaptive Gradient Descent without Descent
Yura Malitsky, Konstantin Mishchenko
We present a strikingly simple proof that two rules are sufficient to automate gradient descent: 1) don't increase the stepsize too fast and 2) don't overstep the local curvature.…
MISO is Making a Comeback With Better Proofs and Rates
Xun Qian, Alibek Sailanbayev, Konstantin Mishchenko +1
MISO, also known as Finito, was one of the first stochastic variance reduced methods discovered, yet its popularity is fairly low. Its initial analysis was significantly limited by…
Revisiting Stochastic Extragradient
Konstantin Mishchenko, Dmitry Kovalev, Egor Shulgin +2
We fix a fundamental issue in the stochastic extragradient method by providing a new sampling strategy that is motivated by approximating implicit updates. Since the existing stoch…
Stochastic Distributed Learning with Gradient Quantization and Variance Reduction
Samuel Horváth, Dmitry Kovalev, Konstantin Mishchenko +2
We consider distributed optimization where the objective function is spread among different devices, each sending incremental model updates to a central server. To alleviate the co…