Accelerating Natural Gradient with Higher-Order Invariance
arXiv:1803.01273
Abstract
An appealing property of the natural gradient is that it is invariant to arbitrary differentiable reparameterizations of the model. However, this invariance property requires infinitesimal steps and is lost in practical implementations with small but finite step sizes. In this paper, we study invariance properties from a combined perspective of Riemannian geometry and numerical differential equation solving. We define the order of invariance of a numerical method to be its convergence order to an invariant solution. We propose to use higher-order integrators and geodesic corrections to obtain more invariant optimization trajectories. We prove the numerical convergence properties of geodesic corrected updates and show that they can be as computationally efficient as plain natural gradient. Experimentally, we demonstrate that invariance leads to faster optimization and our techniques improve on traditional natural gradient in deep neural network training and natural policy gradient for reinforcement learning.
ICML 2018
References in corpus (5)
- Very Deep Convolutional Networks for Large-Scale Image Recognition
- Weight Normalization: A Simple Reparameterization to Accelerate Training of Deep Neural Networks
- Scalable trust-region method for deep reinforcement learning using Kronecker-factored approximation
- Identifying and attacking the saddle point problem in high-dimensional non-convex optimization
- Geodesic acceleration and the small-curvature approximation for nonlinear least squares
Cited by in corpus (7)
- Analysis of Stochastic Gradient Descent in Continuous Time
- A Coordinate-Free Construction of Scalable Natural Gradient
- Handling the Positive-Definite Constraint in the Bayesian Learning Rule
- Quantum Natural Gradient with Geodesic Corrections for Small Shallow Quantum Circuits
- Sinkhorn Natural Gradient for Generative Models
- Noether's Learning Dynamics: Role of Symmetry Breaking in Neural Networks
- Depth Without the Magic: Inductive Bias of Natural Gradient Descent