activity
20122021
most citedA simpler approach to obtaining an O(1/t) convergence rate for the projected stochastic subgradient method

33 citations · 113 across the 8 of their papers we have counts for

collaborators
Showing cs.LGShow all

16 papers · 1 filter

cs.LG2021

SVRG Meets AdaGrad: Painless Variance Reduction

Benjamin Dubois-Taine, Sharan Vaswani, Reza Babanezhad +2

Variance reduction (VR) methods for finite-sum minimization typically require the knowledge of problem-dependent constants that are often unknown and difficult to estimate. To addr…

cs.LG2020

Robust Asymmetric Learning in POMDPs

Andrew Warrington, J. Wilder Lavington, Adam Ścibior +2

Policies for partially observed Markov decision processes can be efficiently learned by imitating policies for the corresponding fully observed Markov decision processes. Unfortuna…

cs.LG2020

Variance-Reduced Methods for Machine Learning

Robert M. Gower, Mark Schmidt, Francis Bach +1

Stochastic optimization lies at the heart of machine learning, and its cornerstone is stochastic gradient descent (SGD), a method introduced over 60 years ago. The last 8 years hav…

cs.LG2020

Regret Bounds without Lipschitz Continuity: Online Learning with Relative-Lipschitz Losses

Yihan Zhou, Victor S. Portella, Mark Schmidt +1

In online convex optimization (OCO), Lipschitz continuity of the functions is commonly assumed in order to obtain sublinear regret. Moreover, many algorithms have only logarithmic…

cs.LG2020

Adaptive Gradient Methods Converge Faster with Over-Parameterization (but you should do a line-search)

Sharan Vaswani, Issam Laradji, Frederik Kunstner +3

Adaptive gradient methods are typically used for training over-parameterized models. To better understand their behaviour, we study a simplistic setting -- smooth, convex losses wi…

cs.LG2019

Fast and Furious Convergence: Stochastic Second Order Methods under Interpolation

Si Yi Meng, Sharan Vaswani, Issam Laradji +2

We consider stochastic second-order methods for minimizing smooth and strongly-convex functions under an interpolation condition satisfied by over-parameterized models. Under this…