From the 1 of 21 linked papers with an AI index.
8 papers · 1 filter
The Dual Averaging Power-Prox Method with Application to Heavy-Tail Incremental Gradient
Yuan Gao, Jeremy Rack, Sebastian U. Stich
We study finite-sum composite optimization under two departures from classical stochastic gradient descent theory that are central in practice: incremental gradient access and heav…
Efficient Gradient Methods for Distributed Saddle Problems
Ruichen Luo, Anton Rodomanov, Sebastian U. Stich
The distributed setting for Saddle Problems (SPs) has recently emerged as a framework for various modern applications in machine learning and multiagent systems. Despite its releva…
DADA: Dual Averaging with Distance Adaptation
Mohammad Moshtaghifar, Anton Rodomanov, Daniil Vankov +1
We present a novel universal gradient method for solving convex optimization problems. Our algorithm, Dual Averaging with Distance Adaptation (DADA), is based on the classical sche…
Composite Optimization with Error Feedback: the Dual Averaging Approach
Yuan Gao, Anton Rodomanov, Jeremy Rack +1
Communication efficiency is a central challenge in distributed machine learning training, and message compression is a widely used solution. However, standard Error Feedback (EF) m…
Accelerated Distributed Optimization with Compression and Error Feedback
Yuan Gao, Anton Rodomanov, Jeremy Rack +1
Modern machine learning tasks often involve massive datasets and models, necessitating distributed optimization algorithms with reduced communication overhead. Communication compre…
Optimizing -Smooth Functions by Gradient Methods
Daniil Vankov, Anton Rodomanov, Angelia Nedich +2
We study gradient methods for optimizing -smooth functions, a class that generalizes Lipschitz-smooth functions and has gained attention for its relevance in machine le…