Practical Riemannian Neural Networks
arXiv:1602.08007
Abstract
We provide the first experimental results on non-synthetic datasets for the quasi-diagonal Riemannian gradient descents for neural networks introduced in [Ollivier, 2015]. These include the MNIST, SVHN, and FACE datasets as well as a previously unpublished electroencephalogram dataset. The quasi-diagonal Riemannian algorithms consistently beat simple stochastic gradient gradient descents by a varying margin. The computational overhead with respect to simple backpropagation is around a factor . Perhaps more interestingly, these methods also reach their final performance quickly, thus requiring fewer training epochs and a smaller total computation time. We also present an implementation guide to these Riemannian gradient descents for neural networks, showing how the quasi-diagonal versions can be implemented with minimal effort on top of existing routines which compute gradients.
Cited by in corpus (7)
- Fisher Information and Natural Gradient Learning of Random Deep Networks
- The Extended Kalman Filter is a Natural Gradient Descent in Trajectory Space
- When Does Preconditioning Help or Hurt Generalization?
- True Asymptotic Natural Gradient Optimization
- Diagonal Rescaling For Neural Networks
- Train Feedfoward Neural Network with Layer-wise Adaptive Rate via Approximating Back-matching Propagation
- Channel-Directed Gradients for Optimization of Convolutional Neural Networks