New insights and perspectives on the natural gradient method
arXiv:1412.1193
Abstract
Natural gradient descent is an optimization method traditionally motivated from the perspective of information geometry, and works well for many applications as an alternative to stochastic gradient descent. In this paper we critically analyze this method and its properties, and show how it can be viewed as a type of 2nd-order optimization method, with the Fisher information matrix acting as a substitute for the Hessian. In many important cases, the Fisher information matrix is shown to be equivalent to the Generalized Gauss-Newton matrix, which both approximates the Hessian, but also has certain properties that favor its use over the Hessian. This perspective turns out to have significant implications for the design of a practical and robust natural gradient optimizer, as it motivates the use of techniques like trust regions and Tikhonov regularization. Additionally, we make a series of contributions to the understanding of natural gradient and 2nd-order methods, including: a thorough analysis of the convergence speed of stochastic natural gradient descent (and more general stochastic 2nd-order methods) as applied to convex quadratics, a critical examination of the oft-used "empirical" approximation of the Fisher matrix, and an analysis of the (approximate) parameterization invariance property possessed by natural gradient methods (which we show also holds for certain other curvature, but notably not the Hessian).
Minor corrections from previous version and fixed typos. Official JMLR version
Cited by in corpus (117)
- A continual learning survey: Defying forgetting in classification tasks
- Three scenarios for continual learning
- Optimizing Neural Networks with Kronecker-factored Approximate Curvature
- Lookahead Optimizer: k steps forward, 1 step back
- Generative replay with feedback connections as a general strategy for continual learning
- Optimization for deep learning: theory and algorithms
- On Quadratic Penalties in Elastic Weight Consolidation
- Machine Unlearning: Solutions and Challenges
- Dissecting Adam: The Sign, Magnitude and Variance of Stochastic Gradients
- Critical Learning Periods in Deep Neural Networks
- Fast Convergence of Natural Gradient Descent for Overparameterized Neural Networks
- LOGAN: Latent Optimisation for Generative Adversarial Networks
- Normalization Techniques in Training DNNs: Methodology, Analysis and Application
- Where is the Information in a Deep Neural Network?
- A Survey of Optimization Methods from a Machine Learning Perspective
- Understanding Short-Horizon Bias in Stochastic Meta-Optimization
- A Kronecker-factored approximate Fisher matrix for convolution layers
- Amortized Variational Inference: A Systematic Review
- Laplace Redux -- Effortless Bayesian Deep Learning
- Information geometry under hierarchical quantum measurement
- Automatic Differentiable Monte Carlo: Theory and Application
- A practical tutorial on Variational Bayes
- Time Matters in Regularizing Deep Networks: Weight Decay and Data Augmentation Affect Early Learning Dynamics, Matter Little Near Convergence
- Local Adaptivity in Federated Learning: Convergence and Consistency
- Early Detection of COVID-19 Hotspots Using Spatio-Temporal Data
- Statistical Inference for the Population Landscape via Moment Adjusted Stochastic Gradients
- Generalized Variational Continual Learning
- Metric Gaussian Variational Inference
- Which Algorithmic Choices Matter at Which Batch Sizes? Insights From a Noisy Quadratic Model
- Continual Learning in Neural Networks
- A Study of Gradient Variance in Deep Learning
- Bayesian Deep Learning via Subnetwork Inference
- The Extended Kalman Filter is a Natural Gradient Descent in Trajectory Space
- Certifiable Machine Unlearning for Linear Models
- Visualizing high-dimensional loss landscapes with Hessian directions
- An Empirical Study of Large-Batch Stochastic Gradient Descent with Structured Covariance Noise
- PAC-Bayes Information Bottleneck
- Optimal transport natural gradient for statistical manifolds with continuous sample space
- Improving predictions of Bayesian neural nets via local linearization
- Information-geometry of physics-informed statistical manifolds and its use in data assimilation
- RelatIF: Identifying Explanatory Training Examples via Relative Influence
- Learning Rates as a Function of Batch Size: A Random Matrix Theory Approach to Neural Network Training
- Kalman meets Bellman: Improving Policy Evaluation through Value Tracking
- Delta-STN: Efficient Bilevel Optimization for Neural Networks using Structured Response Jacobians
- Large-Scale Deep Learning Optimizations: A Comprehensive Survey
- Eigenvalue Corrected Noisy Natural Gradient
- Quantum Natural Policy Gradients: Towards Sample-Efficient Reinforcement Learning
- On the interplay between noise and curvature and its effect on optimization and generalization
- Biological credit assignment through dynamic inversion of feedforward networks
- Aggregated Momentum: Stability Through Passive Damping
- A Coordinate-Free Construction of Scalable Natural Gradient
- The Bayesian Learning Rule
- True Asymptotic Natural Gradient Optimization
- Solving general elliptical mixture models through an approximate Wasserstein manifold
- Dynamics and Reachability of Learning Tasks
- GradSign: Model Performance Inference with Theoretical Insights
- Diagonal Rescaling For Neural Networks
- On Kalman-Bucy filters, linear quadratic control and active inference
- Second-order optimisation strategies for neural network quantum states
- Sketchy Empirical Natural Gradient Methods for Deep Learning
- Trust Region Value Optimization using Kalman Filtering
- Analytic natural gradient updates for Cholesky factor in Gaussian variational approximation
- Unifying Regularisation Methods for Continual Learning
- On the Acceleration of L-BFGS with Second-Order Information and Stochastic Batches
- Resource-constrained Federated Edge Learning with Heterogeneous Data: Formulation and Analysis
- Training Neural Networks with Fixed Sparse Masks
- M-FAC: Efficient Matrix-Free Approximations of Second-Order Information
- Natural continual learning: success is a journey, not (just) a destination
- Posterior Meta-Replay for Continual Learning
- Fast Approximation of the Gauss-Newton Hessian Matrix for the Multilayer Perceptron
- Evaluating the Implicit Midpoint Integrator for Riemannian Manifold Hamiltonian Monte Carlo
- Critical Learning Periods in Federated Learning
- Learnable Uncertainty under Laplace Approximations
- Incremental Learning for Personalized Recommender Systems
- Second-Order Neural ODE Optimizer
- Disentangling the Gauss-Newton Method and Approximate Inference for Neural Networks
- Sampling with Mirrored Stein Operators
- Reweighting Augmented Samples by Minimizing the Maximal Expected Loss
- Convergence of Online Adaptive and Recurrent Optimization Algorithms
- Block Mean Approximation for Efficient Second Order Optimization
- TDprop: Does Jacobi Preconditioning Help Temporal Difference Learning?
- Meta-Learning with Hessian-Free Approach in Deep Neural Nets Training
- Model-agnostic out-of-distribution detection using combined statistical tests
- A Neural Network model with Bidirectional Whitening
- Adaptive Smoothing Path Integral Control
- Stochastic natural gradient descent draws posterior samples in function space
- The Information Geometry of Unsupervised Reinforcement Learning
- Cauchy noise loss for stochastic optimization of random matrix models via free deterministic equivalents
- Noether's Learning Dynamics: Role of Symmetry Breaking in Neural Networks
- ViViT: Curvature access through the generalized Gauss-Newton's low-rank structure
- Multi-resolution spatio-temporal prediction with application to wind power generation
- DIRA: Dynamic Domain Incremental Regularised Adaptation
- Rethinking Gauss-Newton for learning over-parameterized models
- Beyond Folklore: A Scaling Calculus for the Design and Initialization of ReLU Networks
- Self-Tuning Stochastic Optimization with Curvature-Aware Gradient Filtering
- A Trace-restricted Kronecker-Factored Approximation to Natural Gradient
- Two-Level K-FAC Preconditioning for Deep Learning
- An iterative K-FAC algorithm for Deep Learning
- Second-order Information in First-order Optimization Methods
- Appearance of Random Matrix Theory in Deep Learning
- Vanishing Curvature and the Power of Adaptive Methods in Randomly Initialized Deep Networks
- On the Locality of the Natural Gradient for Deep Learning
- Understanding Approximate Fisher Information for Fast Convergence of Natural Gradient Descent in Wide Neural Networks
- Knowledge-Adaptation Priors
- Discriminative Bayesian filtering lends momentum to the stochastic Newton method for minimizing log-convex functions
- Hierarchical model-based policy optimization: from actions to action sequences and back
- Depth Without the Magic: Inductive Bias of Natural Gradient Descent
- TENGraD: Time-Efficient Natural Gradient Descent with Exact Fisher-Block Inversion
- Mirror-Descent Inverse Kinematics for Box-constrained Joint Space
- Pulling back information geometry
- Tensor Normal Training for Deep Learning Models
- The Brownian motion in the transformer model
- Approximate Bayesian Optimisation for Neural Networks
- Adaptive norms for deep learning with regularized Newton methods
- Avoiding Inference Heuristics in Few-shot Prompt-based Finetuning
- Practical Bayesian Learning of Neural Networks via Adaptive Optimisation Methods
- On the Variance of the Fisher Information for Deep Learning