32 citations · 38 across the 2 of their papers we have counts for
5 papers
Which Algorithmic Choices Matter at Which Batch Sizes? Insights From a Noisy Quadratic Model
Guodong Zhang, Lala Li, Zachary Nado +5
Increasing the batch size is a popular way to speed up neural network training, but beyond some critical batch size, larger batch sizes yield diminishing returns. In this work, we…
Adversarial Robustness through Local Linearization
Chongli Qin, James Martens, Sven Gowal +6
Adversarial training is an effective methodology for training deep neural networks that are robust against adversarial, norm-bounded perturbations. However, the computational cost…
Differentiable Game Mechanics
Alistair Letcher, David Balduzzi, Sebastien Racaniere +4
Deep learning is built on the foundational guarantee that gradient descent on an objective function converges to local minima. Unfortunately, this guarantee fails in settings, such…
Fast Convergence of Natural Gradient Descent for Overparameterized Neural Networks
Guodong Zhang, James Martens, Roger Grosse
Natural gradient descent has proven effective at mitigating the effects of pathological curvature in neural network optimization, but little is known theoretically about its conver…
On the Variance of Unbiased Online Recurrent Optimization
Tim Cooijmans, James Martens
The recently proposed Unbiased Online Recurrent Optimization algorithm (UORO, arXiv:1702.05043) uses an unbiased approximation of RTRL to achieve fully online gradient-based learni…