Fast Approximate Natural Gradient Descent in a Kronecker-factored Eigenbasis
arXiv:1806.03884
Abstract
Optimization algorithms that leverage gradient covariance information, such as variants of natural gradient descent (Amari, 1998), offer the prospect of yielding more effective descent directions. For models with many parameters, the covariance matrix they are based on becomes gigantic, making them inapplicable in their original form. This has motivated research into both simple diagonal approximations and more sophisticated factored approximations such as KFAC (Heskes, 2000; Martens & Grosse, 2015; Grosse & Martens, 2016). In the present work we draw inspiration from both to propose a novel approximation that is provably better than KFAC and amendable to cheap partial updates. It consists in tracking a diagonal variance, not in parameter coordinates, but in a Kronecker-factored eigenbasis, in which the diagonal approximation is likely to be more effective. Experiments show improvements over KFAC in optimization speed for several deep network architectures.
Cited by in corpus (24)
- Practical Quasi-Newton Methods for Training Deep Neural Networks
- Descending through a Crowded Valley - Benchmarking Deep Learning Optimizers
- Scalable Second Order Optimization for Deep Learning
- Gram-Gauss-Newton Method: Learning Overparameterized Neural Networks for Regression Problems
- Which Algorithmic Choices Matter at Which Batch Sizes? Insights From a Noisy Quadratic Model
- Discretizing Continuous Action Space for On-Policy Optimization
- Dissecting Hessian: Understanding Common Structure of Hessian in Neural Networks
- Estimating Model Uncertainty of Neural Networks in Sparse Information Form
- KAISA: An Adaptive Second-Order Optimizer Framework for Deep Neural Networks
- Optimization of Graph Neural Networks with Natural Gradient Descent
- Eigenvalue Corrected Noisy Natural Gradient
- On the interplay between noise and curvature and its effect on optimization and generalization
- Kernelized Wasserstein Natural Gradient
- DDPNOpt: Differential Dynamic Programming Neural Optimizer
- Continual Learning with Extended Kronecker-factored Approximate Curvature
- A Differential Game Theoretic Neural Optimizer for Training Residual Networks
- Second-Order Neural ODE Optimizer
- A Trace-restricted Kronecker-Factored Approximation to Natural Gradient
- Robust Out-of-Distribution Detection on Deep Probabilistic Generative Models
- Two-Level K-FAC Preconditioning for Deep Learning
- Dynamic Game Theoretic Neural Optimizer
- Accelerating Distributed K-FAC with Smart Parallelism of Computing and Communication Tasks
- TENGraD: Time-Efficient Natural Gradient Descent with Exact Fisher-Block Inversion
- A Novel Structured Natural Gradient Descent for Deep Learning