33 citations · 113 across the 8 of their papers we have counts for
16 papers · 1 filter
SVRG Meets AdaGrad: Painless Variance Reduction
Benjamin Dubois-Taine, Sharan Vaswani, Reza Babanezhad +2
Variance reduction (VR) methods for finite-sum minimization typically require the knowledge of problem-dependent constants that are often unknown and difficult to estimate. To addr…
Robust Asymmetric Learning in POMDPs
Andrew Warrington, J. Wilder Lavington, Adam Ścibior +2
Policies for partially observed Markov decision processes can be efficiently learned by imitating policies for the corresponding fully observed Markov decision processes. Unfortuna…
Variance-Reduced Methods for Machine Learning
Robert M. Gower, Mark Schmidt, Francis Bach +1
Stochastic optimization lies at the heart of machine learning, and its cornerstone is stochastic gradient descent (SGD), a method introduced over 60 years ago. The last 8 years hav…
Regret Bounds without Lipschitz Continuity: Online Learning with Relative-Lipschitz Losses
Yihan Zhou, Victor S. Portella, Mark Schmidt +1
In online convex optimization (OCO), Lipschitz continuity of the functions is commonly assumed in order to obtain sublinear regret. Moreover, many algorithms have only logarithmic…
Adaptive Gradient Methods Converge Faster with Over-Parameterization (but you should do a line-search)
Sharan Vaswani, Issam Laradji, Frederik Kunstner +3
Adaptive gradient methods are typically used for training over-parameterized models. To better understand their behaviour, we study a simplistic setting -- smooth, convex losses wi…
Fast and Furious Convergence: Stochastic Second Order Methods under Interpolation
Si Yi Meng, Sharan Vaswani, Issam Laradji +2
We consider stochastic second-order methods for minimizing smooth and strongly-convex functions under an interpolation condition satisfied by over-parameterized models. Under this…