81 citations · 159 across the 22 of their papers we have counts for
Showing 2025Show all
3 papers · 1 filter
math.OC2025
Muon is Provably Faster with Momentum Variance Reduction
Xun Qian, Hussein Rammal, Dmitry Kovalev +1
Recent empirical research has demonstrated that deep learning optimizers based on the linear minimization oracle (LMO) over specifically chosen Non-Euclidean norm balls, such as Mu…
cs.LG2025
Understanding Gradient Orthogonalization for Deep Learning via Non-Euclidean Trust-Region Optimization
Dmitry Kovalev
Optimization with matrix gradient orthogonalization has recently demonstrated impressive results in the training of deep neural networks (Jordan et al., 2024; Liu et al., 2025). In…
math.OC2025
On Solving Minimization and Min-Max Problems by First-Order Methods with Relative Error in Gradients
Artem Vasin, Valery Krivchenko, Dmitry Kovalev +6
First-order methods for minimization and saddle point (min-max) problems are widely used for solving large-scale problems, in particular arising in machine learning. The majority o…