4 papers
Glocal Smoothness: Line search and adaptive step sizes can help in theory too!
Curtis Fox, Aaron Mishkin, Sharan Vaswani +1
Iteration complexities for optimizing smooth functions with first-order algorithms are typically stated in terms of a global Lipschitz constant of the gradient, and near-optimal re…
Exploring the loss landscape of regularized neural networks via convex duality
Sungyoon Kim, Aaron Mishkin, Mert Pilanci
We discuss several aspects of the loss landscape of regularized neural networks: the structure of stationary points, connectivity of optimal solutions, path with nonincreasing loss…
Fast Convex Optimization for Two-Layer ReLU Networks: Equivalent Model Classes and Cone Decompositions
Aaron Mishkin, Arda Sahiner, Mert Pilanci
We develop fast algorithms and robust software for convex optimization of two-layer neural networks with ReLU activation functions. Our work leverages a convex reformulation of the…
Faster Convergence of Stochastic Accelerated Gradient Descent under Interpolation
Aaron Mishkin, Mert Pilanci, Mark Schmidt
We prove new convergence rates for a generalized version of stochastic Nesterov acceleration under interpolation conditions. Unlike previous analyses, our approach accelerates any…