81 citations · 116 across the 8 of their papers we have counts for
16 papers
Server-Side Stepsizes and Sampling Without Replacement Provably Help in Federated Optimization
Grigory Malinovsky, Konstantin Mishchenko, Peter Richtárik
We present a theoretical study of server-side optimization in federated learning. Our results are the first to show that the widely popular heuristic of scaling the client updates…
On Seven Fundamental Optimization Challenges in Machine Learning
Konstantin Mishchenko
Many recent successes of machine learning went hand in hand with advances in optimization. The exchange of ideas between these fields has worked both ways, with machine learning bu…
Proximal and Federated Random Reshuffling
Konstantin Mishchenko, Ahmed Khaled, Peter Richtárik
Random Reshuffling (RR), also known as Stochastic Gradient Descent (SGD) without replacement, is a popular and theoretically grounded method for finite-sum minimization. We propose…
Random Reshuffling: Simple Analysis with Vast Improvements
Konstantin Mishchenko, Ahmed Khaled, Peter Richtárik
Random Reshuffling (RR) is an algorithm for minimizing finite-sum functions that utilizes iterative gradient descent steps in conjunction with data reshuffling. Often contrasted wi…
Stochastic Newton and Cubic Newton Methods with Simple Local Linear-Quadratic Rates
Dmitry Kovalev, Konstantin Mishchenko, Peter Richtárik
We present two new remarkably simple stochastic second-order methods for minimizing the average of a very large number of sufficiently smooth and strongly convex functions. The fir…
Adaptive Gradient Descent without Descent
Yura Malitsky, Konstantin Mishchenko
We present a strikingly simple proof that two rules are sufficient to automate gradient descent: 1) don't increase the stepsize too fast and 2) don't overstep the local curvature.…