5 papers
Muon Does Not Converge on Convex Lipschitz Functions
Tetiana Parshakova, Ahmed Khaled, Michael Crawshaw +2
Muon and its variants have shown strong empirical performance in a variety of deep learning tasks. Existing convergence analyses of Muon rely on smoothness assumptions, though argu…
Understanding Outer Optimizers in Local SGD: Learning Rates, Momentum, and Acceleration
Ahmed Khaled, Satyen Kale, Arthur Douillard +3
Modern machine learning often requires training with large batch size, distributed data, and massively parallel compute hardware (like mobile and other edge devices or distributed…
MuonBP: Faster Muon via Block-Periodic Orthogonalization
Ahmed Khaled, Kaan Ozkara, Tao Yu +2
Gradient orthogonalization is a simple strategy that shows great utility in speeding up gradient descent. The Muon optimizer (Jordan, Jin, et al., 2024) combines gradient orthogona…
A Novel Unified Parametric Assumption for Nonconvex Optimization
Artem Riabinin, Ahmed Khaled, Peter Richtárik
Nonconvex optimization is central to modern machine learning, but the general framework of nonconvex optimization yields weak convergence guarantees that are too pessimistic compar…
Directional Smoothness and Gradient Methods: Convergence and Adaptivity
Aaron Mishkin, Ahmed Khaled, Yuanhao Wang +2
We develop new sub-optimality bounds for gradient descent (GD) that depend on the conditioning of the objective along the path of optimization rather than on global, worst-case con…