A Kronecker-factored approximate Fisher matrix for convolution layers
arXiv:1602.01407
Abstract
Second-order optimization methods such as natural gradient descent have the potential to speed up training of neural networks by correcting for the curvature of the loss function. Unfortunately, the exact natural gradient is impractical to compute for large models, and most approximations either require an expensive iterative procedure or make crude approximations to the curvature. We present Kronecker Factors for Convolution (KFC), a tractable approximation to the Fisher matrix for convolutional networks based on a structured probabilistic model for the distribution over backpropagated derivatives. Similarly to the recently proposed Kronecker-Factored Approximate Curvature (K-FAC), each block of the approximate Fisher matrix decomposes as the Kronecker product of small matrices, allowing for efficient inversion. KFC captures important curvature information while still yielding comparably efficient updates to stochastic gradient descent (SGD). We show that the updates are invariant to commonly used reparameterizations, such as centering of the activations. In our experiments, approximate natural gradient descent with KFC was able to train convolutional networks several times faster than carefully tuned SGD. Furthermore, it was able to train the networks in 10-20 times fewer iterations than SGD, suggesting its potential applicability in a distributed setting.
References in corpus (8)
- Sequence to Sequence Learning with Neural Networks
- ADADELTA: An Adaptive Learning Rate Method
- On the difficulty of training Recurrent Neural Networks
- Optimizing Neural Networks with Kronecker-factored Approximate Curvature
- New insights and perspectives on the natural gradient method
- Parallel training of DNNs with Natural Gradient and Parameter Averaging
- On the saddle point problem for non-convex optimization
- Natural Neural Networks
Cited by in corpus (7)
- Riemannian Walk for Incremental Learning: Understanding Forgetting and Intransigence
- Block-diagonal Hessian-free Optimization for Training Neural Networks
- KAISA: An Adaptive Second-Order Optimizer Framework for Deep Neural Networks
- A Coordinate-Free Construction of Scalable Natural Gradient
- Scalable Natural Gradient Langevin Dynamics in Practice
- Train Feedfoward Neural Network with Layer-wise Adaptive Rate via Approximating Back-matching Propagation
- First-Order Preconditioning via Hypergradient Descent