Revisiting Natural Gradient for Deep Networks
arXiv:1301.3584
Abstract
We evaluate natural gradient, an algorithm originally proposed in Amari (1997), for learning deep models. The contributions of this paper are as follows. We show the connection between natural gradient and three other recently proposed methods for training deep models: Hessian-Free (Martens, 2010), Krylov Subspace Descent (Vinyals and Povey, 2012) and TONGA (Le Roux et al., 2008). We describe how one can use unlabeled data to improve the generalization error obtained by natural gradient and empirically evaluate the robustness of the algorithm to the ordering of the training set compared to stochastic gradient descent. Finally we extend natural gradient to incorporate second order information alongside the manifold information and provide a benchmark of the new algorithm using a truncated Newton approach for inverting the metric matrix instead of using a diagonal approximation of it.
Cited by in corpus (23)
- Overcoming catastrophic forgetting in neural networks
- A continual learning survey: Defying forgetting in classification tasks
- Riemannian Walk for Incremental Learning: Understanding Forgetting and Intransigence
- How to Construct Deep Recurrent Neural Networks
- Deep Learning of Representations: Looking Forward
- Pseudo-Rehearsal: Achieving Deep Reinforcement Learning without Catastrophic Forgetting
- A Bayesian Federated Learning Framework with Online Laplace Approximation
- Overcoming Long-term Catastrophic Forgetting through Adversarial Neural Pruning and Synaptic Consolidation
- Strong error analysis for stochastic gradient descent optimization algorithms
- FedLGA: Towards System-Heterogeneity of Federated Learning via Local Gradient Approximation
- Deep Latent Dirichlet Allocation with Topic-Layer-Adaptive Stochastic Gradient Riemannian MCMC
- Automatic Differentiable Monte Carlo: Theory and Application
- DeepStability: A Study of Unstable Numerical Methods and Their Solutions in Deep Learning
- Understanding symmetries in deep networks
- Interstellar: Searching Recurrent Architecture for Knowledge Graph Embedding
- Symmetry-invariant optimization in deep networks
- Second-order optimisation strategies for neural network quantum states
- Posterior Meta-Replay for Continual Learning
- Learned-Norm Pooling for Deep Feedforward and Recurrent Neural Networks
- Relative Natural Gradient for Learning Large Complex Models
- Task-agnostic Continual Learning with Hybrid Probabilistic Models
- A Neural Network model with Bidirectional Whitening
- Weighted Ensemble Models Are Strong Continual Learners