Showing cs.LGShow all
2 papers · 1 filter
cs.LG2020
Adaptive Gradient Methods Converge Faster with Over-Parameterization (but you should do a line-search)
Sharan Vaswani, Issam Laradji, Frederik Kunstner +3
Adaptive gradient methods are typically used for training over-parameterized models. To better understand their behaviour, we study a simplistic setting -- smooth, convex losses wi…
cs.LG2019
Fast and Furious Convergence: Stochastic Second Order Methods under Interpolation
Si Yi Meng, Sharan Vaswani, Issam Laradji +2
We consider stochastic second-order methods for minimizing smooth and strongly-convex functions under an interpolation condition satisfied by over-parameterized models. Under this…