Path-SGD: Path-Normalized Optimization in Deep Neural Networks
arXiv:1506.02617
Abstract
We revisit the choice of SGD for training deep neural networks by reconsidering the appropriate geometry in which to optimize the weights. We argue for a geometry invariant to rescaling of weights that does not affect the output of the network, and suggest Path-SGD, which is an approximate steepest descent method with respect to a path-wise regularizer related to max-norm regularization. Path-SGD is easy and efficient to implement and leads to empirical gains over SGD and AdaGrad.
12 pages, 5 figures
References in corpus (1)
Cited by in corpus (7)
- Layer Normalization
- Exploring Generalization in Deep Learning
- Towards Understanding Generalization of Deep Learning: Perspective of Loss Landscapes
- Implicit Regularization in Deep Learning
- Riemannian approach to batch normalization
- The Implicit Bias of AdaGrad on Separable Data
- Orthogonal and Idempotent Transformations for Learning Deep Neural Networks