500 citations · 1.3k across the 50 of their papers we have counts for
35 papers · 1 filter
Improved Stein Variational Gradient Descent with Importance Weights
Lukang Sun, Peter Richtárik
Stein Variational Gradient Descent (SVGD) is a popular sampling algorithm used in various machine learning tasks. It is well known that SVGD arises from a discretization of the ker…
Adaptive Compression for Communication-Efficient Distributed Training
Maksim Makarenko, Elnur Gasanov, Rustem Islamov +2
We propose Adaptive Compressed Gradient Descent (AdaCGD) - a novel optimization algorithm for communication-efficient training of supervised machine learning models with adaptive c…
Federated Random Reshuffling with Compression and Variance Reduction
Grigory Malinovsky, Peter Richtárik
Random Reshuffling (RR), which is a variant of Stochastic Gradient Descent (SGD) employing sampling without replacement, is an immensely popular method for training supervised mach…
Permutation Compressors for Provably Faster Distributed Nonconvex Optimization
Rafał Szlendak, Alexander Tyurin, Peter Richtárik
We study the MARINA method of Gorbunov et al (2021) -- the current state-of-the-art distributed non-convex optimization method in terms of theoretical communication complexity. The…
Doubly Adaptive Scaled Algorithm for Machine Learning Using Second-Order Information
Majid Jahani, Sergey Rusakov, Zheng Shi +3
We present a novel adaptive optimization algorithm for large-scale machine learning problems. Equipped with a low-cost estimate of local curvature and Lipschitz smoothness, our met…
FedPAGE: A Fast Local Stochastic Gradient Method for Communication-Efficient Federated Learning
Haoyu Zhao, Zhize Li, Peter Richtárik
Federated Averaging (FedAvg, also known as Local-SGD) (McMahan et al., 2017) is a classical federated learning algorithm in which clients run multiple local SGD steps before commun…