12 papers · 1 filter
Convergence Analysis of Muon-type Methods with Inexact LMO in the Degenerate Case
Xun Qian, Peter Richtárik
Muon-type methods have demonstrated potentially superior performance over Adam and its variants, and have shown hyperparameter transferability across model sizes when specific norm…
A Unified Primal-Dual Recipe for Accelerating Three-Operator Splitting Methods
Abdurakhmon Sadiev, Laurent Condat, Peter Richtárik
Composite optimization problems, formulated as the minimization of three functions, are ubiquitous in large-scale machine learning and signal processing. While state-of-the-art spl…
Rennala MVR: Improved Time Complexity for Parallel Stochastic Optimization via Momentum-Based Variance Reduction
Zhirayr Tovmasyan, Artavazd Maranjyan, Peter Richtárik
Large-scale machine learning models are trained on clusters of machines that exhibit heterogeneous performance due to hardware variability, network delays, and system-level instabi…
Local LMO: Constrained Gradient Optimization via a Local Linear Minimization Oracle
Peter Richtárik, Kaja Gruntkowska, Hanmin Li
We design Local LMO - a new projection-free gradient-type method for constrained optimization. The key algorithmic idea is to replace the global linear minimization oracle over the…
Broximal Alignment for Global Non-Convex Optimization
Kaja Gruntkowska, Hanmin Li, Xun Qian +1
Most non-convex optimization theory is built around gradient dynamics, leaving global convergence largely unexplored. The dominant paradigm focuses on stationarity, certifying only…
A Nesterov-Accelerated Primal-Dual Splitting Algorithm for Convex Nonsmooth Optimization
Laurent Condat, Abdurakhmon Sadiev, Peter Richtárik
We investigate the integration of Nesterov-type acceleration into primal-dual methods for structured convex optimization. While proximal splitting algorithms efficiently handle com…