12 papers
On the Limits of Biased Derivative Information for Nonconvex Stochastic Optimization
Anant Shyam, Brian Bullins
We consider the problem of finding -stationary points for , i.e., such that , for smooth, non-convex objectives, where the d…
Can Entry-Wise Clipping Give Spectral Control of Stochastic Gradients?
Zitao Song, Cedar Site Bai, Zhe Zhang +2
Training instabilities such as loss spikes are frequently the result of stochastic gradient noise. Because of rare expressions in language training data, and multiple layer composi…
Decoupling Variance and Scale-Invariant Updates in Adaptive Gradient Descent for Unified Vector and Matrix Optimization
Zitao Song, Cedar Site Bai, Zhe Zhang +2
Adaptive methods like Adam have become the standard for large-scale vector and Euclidean optimization due to their coordinate-wise adaptation with a second-orde…
Mirror-Free Proximal Methods
Abhijeet Vyas, Brian Bullins
We present a \emph{mirror-free} mirror prox (MFMP) algorithm, which extends the classic approach of Nemirovski (2004) to allow for proximal-like updates without the explicit need f…
Beyond First-Order Methods for -Structured Non-Monotone Variational Inequalities
Abhijeet Vyas, Brian Bullins
We propose novel high-order algorithms for a class of -structured non-monotone variational inequalities. In particular, work by Diakonikolas et al. (2021), which introduced…
Online Min-Max Optimization: From Individual Regrets to Cumulative Saddle Points
Abhijeet Vyas, Brian Bullins
We propose and study an online version of min-max optimization based on cumulative saddle points under a variety of performance measures beyond convex-concave settings. After first…