1 citations · 1 across the 6 of their papers we have counts for
3 papers · 1 filter
Glocal Smoothness: Line search and adaptive step sizes can help in theory too!
Curtis Fox, Aaron Mishkin, Sharan Vaswani +1
Iteration complexities for optimizing smooth functions with first-order algorithms are typically stated in terms of a global Lipschitz constant of the gradient, and near-optimal re…
Towards Noise-adaptive, Problem-adaptive (Accelerated) Stochastic Gradient Descent
Sharan Vaswani, Benjamin Dubois-Taine, Reza Babanezhad
We aim to make stochastic gradient descent (SGD) adaptive to (i) the noise in the stochastic gradients and (ii) problem-dependent constants. When minimizing smooth, strongly…
(Accelerated) Noise-adaptive Stochastic Heavy-Ball Momentum
Anh Dang, Reza Babanezhad, Sharan Vaswani
Stochastic heavy ball momentum (SHB) is commonly used to train machine learning models, and often provides empirical improvements over stochastic gradient descent. By primarily foc…