11 papers
Exact Convex Reformulations of Linear Neural Networks via Completely Positive Lifting
Karthik Prakhya, Alp Yurtsever
We show that the training problem of a deep linear neural network under the squared loss admits an exact convex reformulation in a lifted space over a generalized completely positi…
Universal Adaptive Proximal Gradient Methods via Gradient Mapping Accumulation
Zimeng Wang, Alp Yurtsever
We propose an adaptive proximal gradient method for minimizing the sum of two functions, where one is a simple convex function, and the other belongs to one of the three classes: n…
The Multi-Block DC Function Class: Theory, Algorithms, and Applications
Pouria Fatemi, Hoomaan Maskan, Alp Yurtsever +1
We present the Multi-Block DC (BDC) class, a rich class of structured nonconvex functions that admit a DC ("difference-of-convex") decomposition across parameter blocks. This multi…
Generalized Stochastic Gradient Descent with Momentum Methods for Smooth Optimization
Zimeng Wang, Alp Yurtsever
Stochastic gradient descent with momentum (SGDM) methods have become fundamental optimization tools in machine learning, combining the computational efficiency of stochastic gradie…
Revisiting Frank-Wolfe for Structured Nonconvex Optimization
Hoomaan Maskan, Yikun Hou, Suvrit Sra +1
We introduce a new projection-free (Frank-Wolfe) method for optimizing structured nonconvex functions that are expressed as a difference of two convex functions. This problem class…
Randomized Block Coordinate DC Programming
Hoomaan Maskan, Paniz Halvachi, Suvrit Sra +1
We introduce an extension of the Difference of Convex Algorithm (DCA) in the form of a randomized block coordinate approach for problems with separable structure. For coordinat…