2 citations · 2 across the 8 of their papers we have counts for
9 papers · 1 filter
Free Heavy-Tailed Lunch for Muon: A Theoretical Justification of Empirical Success
Florian Hübler, Thomas Pethick, Suvrit Sra
Non-Euclidean optimisation methods with matrix-valued updates, such as Muon and Scion, have recently shown strong empirical performance for training Transformer models, yet their t…
The Multi-Block DC Function Class: Theory, Algorithms, and Applications
Pouria Fatemi, Hoomaan Maskan, Alp Yurtsever +1
We present the Multi-Block DC (BDC) class, a rich class of structured nonconvex functions that admit a DC ("difference-of-convex") decomposition across parameter blocks. This multi…
Revisiting Frank-Wolfe for Structured Nonconvex Optimization
Hoomaan Maskan, Yikun Hou, Suvrit Sra +1
We introduce a new projection-free (Frank-Wolfe) method for optimizing structured nonconvex functions that are expressed as a difference of two convex functions. This problem class…
Randomized Block Coordinate DC Programming
Hoomaan Maskan, Paniz Halvachi, Suvrit Sra +1
We introduce an extension of the Difference of Convex Algorithm (DCA) in the form of a randomized block coordinate approach for problems with separable structure. For coordinat…
Improved Rates for Stochastic Variance-Reduced Difference-of-Convex Algorithms
Anh Duc Nguyen, Alp Yurtsever, Suvrit Sra +1
In this work, we propose and analyze DCA-PAGE, a novel algorithm that integrates the difference-of-convex algorithm (DCA) with the ProbAbilistic Gradient Estimator (PAGE) to solve…
Linearly Convergent Algorithms for Nonsmooth Problems with Unknown Smooth Pieces
Zhe Zhang, Suvrit Sra
We develop efficient algorithms for optimizing piecewise smooth (PWS) functions where the underlying partition of the domain into smooth pieces is \emph{unknown}. For PWS functions…