3 papers
math.OC2026
Rennala MVR: Improved Time Complexity for Parallel Stochastic Optimization via Momentum-Based Variance Reduction
Zhirayr Tovmasyan, Artavazd Maranjyan, Peter Richtárik
Large-scale machine learning models are trained on clusters of machines that exhibit heterogeneous performance due to hardware variability, network delays, and system-level instabi…
cs.LG2025
Error Feedback for Muon and Friends
Kaja Gruntkowska, Alexander Gaponov, Zhirayr Tovmasyan +1
Recent optimizers like Muon, Scion, and Gluon have pushed the frontier of large-scale deep learning by exploiting layer-wise linear minimization oracles (LMOs) over non-Euclidean n…
math.OC2025
Revisiting Stochastic Proximal Point Methods: Generalized Smoothness and Similarity
Zhirayr Tovmasyan, Grigory Malinovsky, Laurent Condat +1
The growing prevalence of nonsmooth optimization problems in machine learning has spurred significant interest in generalized smoothness assumptions. Among these, the (L0, L1)-smoo…