activity
20172025
most citedDecentralized Stochastic Optimization and Gossip Algorithms with Compressed Communication

185 citations · 707 across the 20 of their papers we have counts for

collaborators
Showing math.OCShow all

7 papers · 1 filter

math.OC2025

Monotone and nonmonotone linearized block coordinate descent methods for nonsmooth composite optimization problems

Yassine Nabou, Lahcen El Bourkhissi, Sebastian U. Stich +1

In this paper, we introduce both monotone and nonmonotone variants of LiBCoD, a \textbf{Li}nearized \textbf{B}lock \textbf{Co}ordinate \textbf{D}escent method for solving composite…

math.OC2025

Revisiting LocalSGD and SCAFFOLD: Improved Rates and Missing Analysis

Ruichen Luo, Sebastian U Stich, Samuel Horváth +1

LocalSGD and SCAFFOLD are widely used methods in distributed stochastic optimization, with numerous applications in machine learning, large-scale data processing, and federated lea…

math.OC20204 cited

A Linearly Convergent Algorithm for Decentralized Optimization: Sending Less Bits for Free!

Dmitry Kovalev, Anastasia Koloskova, Martin Jaggi +2

Decentralized optimization methods enable on-device training of machine learning models without a central coordinator. In many scenarios communication between devices is energy dem…

math.OC201981 cited

Stochastic Distributed Learning with Gradient Quantization and Variance Reduction

Samuel Horváth, Dmitry Kovalev, Konstantin Mishchenko +2

We consider distributed optimization where the objective function is spread among different devices, each sending incremental model updates to a central server. To alleviate the co…

math.OC2018

Efficient Greedy Coordinate Descent for Composite Problems

Sai Praneeth Karimireddy, Anastasia Koloskova, Sebastian U. Stich +1

Coordinate descent with random coordinate selection is the current state of the art for many large scale optimization problems. However, greedy selection of the steepest coordinate…

math.OC2018

Local SGD Converges Fast and Communicates Little

Sebastian U. Stich

Mini-batch stochastic gradient descent (SGD) is state of the art in large scale distributed training. The scheme can reach a linear speedup with respect to the number of workers, b…