Equilibrated adaptive learning rates for non-convex optimization
arXiv:1502.04390
Abstract
Parameter-specific adaptive learning rate methods are computationally efficient ways to reduce the ill-conditioning problems encountered when training large deep networks. Following recent work that strongly suggests that most of the critical points encountered when training such networks are saddle points, we find how considering the presence of negative eigenvalues of the Hessian could help us design better suited adaptive learning rate schemes. We show that the popular Jacobi preconditioner has undesirable behavior in the presence of both positive and negative curvature, and present theoretical and empirical evidence that the so-called equilibration preconditioner is comparatively better suited to non-convex problems. We introduce a novel adaptive learning rate scheme, called ESGD, based on the equilibration preconditioner. Our experiments show that ESGD performs as well or better than RMSProp in terms of convergence speed, always clearly improving over plain stochastic gradient descent.
References in corpus (4)
Cited by in corpus (22)
- Preconditioned Stochastic Gradient Langevin Dynamics for Deep Neural Networks
- Handwritten Isolated Bangla Compound Character Recognition: a new benchmark using a novel deep learning approach
- Preconditioned Stochastic Gradient Descent
- On the Convergence of A Class of Adam-Type Algorithms for Non-Convex Optimization
- Neural Chinese Named Entity Recognition via CNN-LSTM-CRF and Joint Training with Word Segmentation
- Strong error analysis for stochastic gradient descent optimization algorithms
- Prompt- and Trait Relation-aware Cross-prompt Essay Trait Scoring
- A Deep Generative Deconvolutional Image Model
- Is the PPG signal chaotic?
- Neural Chinese Word Segmentation with Lexicon and Unlabeled Data via Posterior Regularization
- Bayesian Sparse learning with preconditioned stochastic gradient MCMC and its applications
- Robust and efficient algorithms for high-dimensional black-box quantum optimization
- Scalable Balanced Training of Conditional Generative Adversarial Neural Networks on Image Data
- Neural Taylor Approximations: Convergence and Exploration in Rectifier Networks
- An adaptive Hessian approximated stochastic gradient MCMC method
- Gravity Optimizer: a Kinematic Approach on Optimization in Deep Learning
- TDprop: Does Jacobi Preconditioning Help Temporal Difference Learning?
- A Survey on Large-scale Machine Learning
- Single-Solution Hypervolume Maximization and its use for Improving Generalization of Neural Networks
- Primitive Agentic First-Order Optimization
- DEAM: Adaptive Momentum with Discriminative Weight for Stochastic Optimization
- PrecoG: an efficient unitary split preconditioner for the transform-domain LMS filter via graph Laplacian regularization