3 papers
cs.LG2026
Flatland: The Adventures of Gradient Descent with Large Step Sizes
Leonardo Galli, Curtis Fox, Wiebke Bartolomaeus +2
The training of neural networks often entails objective functions that are not globally -smooth. For these functions, it is both theoretically and practically difficult to reply…
math.OC2026
Glocal Smoothness: Line search and adaptive step sizes can help in theory too!
Curtis Fox, Aaron Mishkin, Sharan Vaswani +1
Iteration complexities for optimizing smooth functions with first-order algorithms are typically stated in terms of a global Lipschitz constant of the gradient, and near-optimal re…
math.OC2025
Faster Convergence of Stochastic Accelerated Gradient Descent under Interpolation
Aaron Mishkin, Mert Pilanci, Mark Schmidt
We prove new convergence rates for a generalized version of stochastic Nesterov acceleration under interpolation conditions. Unlike previous analyses, our approach accelerates any…