Optimizing Neural Networks with Kronecker-factored Approximate Curvature
arXiv:1503.05671
Abstract
We propose an efficient method for approximating natural gradient descent in neural networks which we call Kronecker-Factored Approximate Curvature (K-FAC). K-FAC is based on an efficiently invertible approximation of a neural network's Fisher information matrix which is neither diagonal nor low-rank, and in some cases is completely non-sparse. It is derived by approximating various large blocks of the Fisher (corresponding to entire layers) as being the Kronecker product of two much smaller matrices. While only several times more expensive to compute than the plain stochastic gradient, the updates produced by K-FAC make much more progress optimizing the objective, which results in an algorithm that can be much faster than stochastic gradient descent with momentum in practice. And unlike some previously proposed approximate natural-gradient/Newton methods which use high-quality non-diagonal curvature matrices (such as Hessian-free optimization), K-FAC works very well in highly stochastic optimization regimes. This is because the cost of storing and inverting K-FAC's approximation to the curvature matrix does not depend on the amount of data used to estimate it, which is a feature typically associated only with diagonal or low-rank approximations to the curvature matrix.
Reduction ratio formula corrected. Removed incorrect claim about geodesics in footnote
References in corpus (5)
Cited by in corpus (190)
- Riemannian Walk for Incremental Learning: Understanding Forgetting and Intransigence
- Weight Normalization: A Simple Reparameterization to Accelerate Training of Deep Neural Networks
- Don't Decay the Learning Rate, Increase the Batch Size
- Scalable trust-region method for deep reinforcement learning using Kronecker-factored approximation
- A DIRT-T Approach to Unsupervised Domain Adaptation
- Lookahead Optimizer: k steps forward, 1 step back
- Learning to learn by gradient descent by gradient descent
- Composing graphical models with neural networks for structured representations and fast inference
- On Empirical Comparisons of Optimizers for Deep Learning
- Large Batch Optimization for Deep Learning: Training BERT in 76 minutes
- Bayesian Model-Agnostic Meta-Learning
- Measuring the Effects of Data Parallelism on Neural Network Training
- Autonomous Unmanned Aerial Vehicle Navigation using Reinforcement Learning: A Systematic Review
- Preconditioned Stochastic Gradient Descent
- Fisher-Rao Metric, Geometry, and Complexity of Neural Networks
- Reinforcement Learning Algorithms: An Overview and Classification
- Sharpness-Aware Minimization for Efficiently Improving Generalization
- Neural Policy Gradient Methods: Global Optimality and Rates of Convergence
- Learning values across many orders of magnitude
- Convergence guarantees for RMSProp and ADAM in non-convex optimization and an empirical comparison to Nesterov acceleration
- Pitfalls of In-Domain Uncertainty Estimation and Ensembling in Deep Learning
- Implicit Regularization in Deep Learning
- On the Expressive Power of Deep Neural Networks
- Task Agnostic Continual Learning Using Online Variational Bayes
- Meta-Learning with Warped Gradient Descent
- Critical Learning Periods in Deep Neural Networks
- A Progressive Batching L-BFGS Method for Machine Learning
- Fast Convergence of Natural Gradient Descent for Overparameterized Neural Networks
- On Warm-Starting Neural Network Training
- Shampoo: Preconditioned Stochastic Tensor Optimization
- Normalization Techniques in Training DNNs: Methodology, Analysis and Application
- Three Mechanisms of Weight Decay Regularization
- BRECQ: Pushing the Limit of Post-Training Quantization by Block Reconstruction
- Universal Statistics of Fisher Information in Deep Neural Networks: Mean Field Approach
- A Survey of Optimization Methods from a Machine Learning Perspective
- Continual Learning and Catastrophic Forgetting
- Understanding Short-Horizon Bias in Stochastic Meta-Optimization
- Stein Variational Gradient Descent With Matrix-Valued Kernels
- Recent Advances of Continual Learning in Computer Vision: An Overview
- Beyond Convexity -- Contraction and Global Convergence of Gradient Descent
- Practical Quasi-Newton Methods for Training Deep Neural Networks
- WoodFisher: Efficient Second-Order Approximation for Neural Network Compression
- Being Bayesian, Even Just a Bit, Fixes Overconfidence in ReLU Networks
- Descending through a Crowded Valley - Benchmarking Deep Learning Optimizers
- Tighter risk certificates for neural networks
- The Break-Even Point on Optimization Trajectories of Deep Neural Networks
- A Theoretical Framework for Target Propagation
- A Kronecker-factored approximate Fisher matrix for convolution layers
- PyHessian: Neural Networks Through the Lens of the Hessian
- Scalable Second Order Optimization for Deep Learning
- Natural Environment Benchmarks for Reinforcement Learning
- Gram-Gauss-Newton Method: Learning Overparameterized Neural Networks for Regression Problems
- Laplace Redux -- Effortless Bayesian Deep Learning
- Trust-PCL: An Off-Policy Trust Region Method for Continuous Control
- Noisy Natural Gradient as Variational Inference
- Automatic Differentiable Monte Carlo: Theory and Application
- Meta-Learning Symmetries by Reparameterization
- Deep Frank-Wolfe For Neural Network Optimization
- Neural Parametric Fokker-Planck Equations
- Boosting Trust Region Policy Optimization by Normalizing Flows Policy
- Large-Scale Distributed Second-Order Optimization Using Kronecker-Factored Approximate Curvature for Deep Convolutional Neural Networks
- Hessian-based toolbox for reliable and interpretable machine learning in physics
- Connecting geometry and performance of two-qubit parameterized quantum circuits
- Measuring and regularizing networks in function space
- On the distance between two neural networks and the stability of learning
- of two-dimensional electron gas: a neural canonical transformation study
- Task Agnostic Continual Learning Using Online Variational Bayes with Fixed-Point Updates
- Training Algorithm Matters for the Performance of Neural Network Potential: A Case Study of Adam and the Kalman Filter Optimizers
- Optimization and Generalization of Regularization-Based Continual Learning: a Loss Approximation Viewpoint
- Which Algorithmic Choices Matter at Which Batch Sizes? Insights From a Noisy Quadratic Model
- Discretizing Continuous Action Space for On-Policy Optimization
- Distributed Learning of Deep Neural Networks using Independent Subnet Training
- Kronecker Determinantal Point Processes
- Riemannian approach to batch normalization
- Robust Frequent Directions with Application in Online Learning
- Inexact Newton Methods for Stochastic Nonconvex Optimization with Applications to Neural Network Training
- Dissecting Hessian: Understanding Common Structure of Hessian in Neural Networks
- Spectral Normalisation for Deep Reinforcement Learning: an Optimisation Perspective
- Scalable Bayesian Meta-Learning through Generalized Implicit Gradients
- An Empirical Study of Large-Batch Stochastic Gradient Descent with Structured Covariance Noise
- The Two Regimes of Deep Network Training
- When Does Preconditioning Help or Hurt Generalization?
- Neural Replicator Dynamics
- Scalable Marginal Likelihood Estimation for Model Selection in Deep Learning
- Block-diagonal Hessian-free Optimization for Training Neural Networks
- Second-order step-size tuning of SGD for non-convex optimization
- Improving predictions of Bayesian neural nets via local linearization
- KAISA: An Adaptive Second-Order Optimizer Framework for Deep Neural Networks
- Eigenvalue Corrected Noisy Natural Gradient
- Learning Rates as a Function of Batch Size: A Random Matrix Theory Approach to Neural Network Training
- Fast Predictive Uncertainty for Classification with Bayesian Deep Networks
- Pathological spectra of the Fisher information metric and its variants in deep neural networks
- Kronecker Recurrent Units
- The Normalization Method for Alleviating Pathological Sharpness in Wide Neural Networks
- Trusting SVM for Piecewise Linear CNNs
- Scalable Natural Gradient Langevin Dynamics in Practice
- Rayleigh-Gauss-Newton optimization with enhanced sampling for variational Monte Carlo
- True Asymptotic Natural Gradient Optimization
- A Coordinate-Free Construction of Scalable Natural Gradient
- Aggregated Momentum: Stability Through Passive Damping
- Batch Normalization Preconditioning for Neural Network Training
- Information Geometry of Orthogonal Initializations and Training
- Information Newton's flow: second-order optimization method in probability space
- Accelerating Natural Gradient with Higher-Order Invariance
- Learning Practically Feasible Policies for Online 3D Bin Packing
- Fast and Furious Convergence: Stochastic Second Order Methods under Interpolation
- Training Neural Networks for and by Interpolation
- Deep learning quantum Monte Carlo for solids
- Second-order optimisation strategies for neural network quantum states
- Diagonal Rescaling For Neural Networks
- Sketchy Empirical Natural Gradient Methods for Deep Learning
- Kernelized Wasserstein Natural Gradient
- Analytic Insights into Structure and Rank of Neural Network Hessian Maps
- Ab-Initio Potential Energy Surfaces by Pairing GNNs with Neural Wave Functions
- On the Promise of the Stochastic Generalized Gauss-Newton Method for Training DNNs
- Natural continual learning: success is a journey, not (just) a destination
- Fast Approximation of the Gauss-Newton Hessian Matrix for the Multilayer Perceptron
- A Generalizable Approach to Learning Optimizers
- M-FAC: Efficient Matrix-Free Approximations of Second-Order Information
- On the Acceleration of L-BFGS with Second-Order Information and Stochastic Batches
- DDPNOpt: Differential Dynamic Programming Neural Optimizer
- Data-Dependent Path Normalization in Neural Networks
- LocoProp: Enhancing BackProp via Local Loss Optimization
- Deep quantum Monte Carlo approach for polaritonic chemistry
- Interactive Label Cleaning with Example-based Explanations
- Learnable Uncertainty under Laplace Approximations
- Multi-Agent Interactions Modeling with Correlated Policies
- The Variational Predictive Natural Gradient
- Inefficiency of K-FAC for Large Batch Size Training
- Synergy between deep neural networks and the variational Monte Carlo method for small clusters
- Does the Data Induce Capacity Control in Deep Learning?
- Neumann Optimizer: A Practical Optimization Algorithm for Deep Neural Networks
- Continual Learning: Tackling Catastrophic Forgetting in Deep Neural Networks with Replay Processes
- Thoughts on the Consistency between Ricci Flow and Neural Network Behavior
- Efficient Implementation of Second-Order Stochastic Approximation Algorithms in High-Dimensional Problems
- Beyond the Mean-Field: Structured Deep Gaussian Processes Improve the Predictive Uncertainties
- What Deep CNNs Benefit from Global Covariance Pooling: An Optimization Perspective
- Error Bounds and Applications for Stochastic Approximation with Non-Decaying Gain
- Reverse engineering learned optimizers reveals known and novel mechanisms
- Layer-wise Conditioning Analysis in Exploring the Learning Dynamics of DNNs
- An Empirical Analysis of Proximal Policy Optimization with Kronecker-factored Natural Gradients
- First-Order Preconditioning via Hypergradient Descent
- Concurrent Adversarial Learning for Large-Batch Training
- Non-Parametric Calibration for Classification
- LAGC: Lazily Aggregated Gradient Coding for Straggler-Tolerant and Communication-Efficient Distributed Learning
- Stochastic natural gradient descent draws posterior samples in function space
- Meta-Learning with Hessian-Free Approach in Deep Neural Nets Training
- Training Neural Networks in Single vs Double Precision
- Model-agnostic out-of-distribution detection using combined statistical tests
- Train Feedfoward Neural Network with Layer-wise Adaptive Rate via Approximating Back-matching Propagation
- Biologically inspired architectures for sample-efficient deep reinforcement learning
- A Neural Network model with Bidirectional Whitening
- Towards Query-Efficient Black-Box Adversary with Zeroth-Order Natural Gradient Descent
- Appearance of Random Matrix Theory in Deep Learning
- Hindsight Experience Replay with Kronecker Product Approximate Curvature
- A Dynamic Sampling Adaptive-SGD Method for Machine Learning
- Improving SGD convergence by online linear regression of gradients in multiple statistically relevant directions
- Online Second Order Methods for Non-Convex Stochastic Optimizations
- Noether's Learning Dynamics: Role of Symmetry Breaking in Neural Networks
- Rethinking Gauss-Newton for learning over-parameterized models
- Modular Block-diagonal Curvature Approximations for Feedforward Architectures
- Dissecting Non-Vacuous Generalization Bounds based on the Mean-Field Approximation
- ViViT: Curvature access through the generalized Gauss-Newton's low-rank structure
- Structured second-order methods via natural gradient descent
- Training Efficiency and Robustness in Deep Learning
- Reinforcement Learning for Robust Missile Autopilot Design
- Accelerated Information Gradient flow
- Tensor Normal Training for Deep Learning Models
- Accelerating Distributed K-FAC with Smart Parallelism of Computing and Communication Tasks
- Rapid training of deep neural networks without skip connections or normalization layers using Deep Kernel Shaping
- Meta-Learning with Variational Bayes
- On the Locality of the Natural Gradient for Deep Learning
- Robust Trust Region for Weakly Supervised Segmentation
- Enhance Curvature Information by Structured Stochastic Quasi-Newton Methods
- Understanding Approximate Fisher Information for Fast Convergence of Natural Gradient Descent in Wide Neural Networks
- Kronecker Factorization for Preventing Catastrophic Forgetting in Large-scale Medical Entity Linking
- Depth Without the Magic: Inductive Bias of Natural Gradient Descent
- Parabolic Approximation Line Search for DNNs
- Preconditioner on Matrix Lie Group for SGD
- Bridging the Gap between Deep Learning and Frustrated Quantum Spin System for Extreme-scale Simulations on New Generation of Sunway Supercomputer
- Bayesian filtering unifies adaptive and non-adaptive neural network optimization methods
- Adaptive norms for deep learning with regularized Newton methods
- Whitening and second order optimization both make information in the dataset unusable during training, and can reduce or prevent generalization
- TENGraD: Time-Efficient Natural Gradient Descent with Exact Fisher-Block Inversion
- Pulling back information geometry
- mL-BFGS: A Momentum-based L-BFGS for Distributed Large-Scale Neural Network Optimization
- Deep Neural Network Learning with Second-Order Optimizers -- a Practical Study with a Stochastic Quasi-Gauss-Newton Method
- Riemannian Laplace approximations for Bayesian neural networks
- KF-LAX: Kronecker-factored curvature estimation for control variate optimization in reinforcement learning
- Hamiltonian Monte-Carlo for Orthogonal Matrices