On the Convergence of Adam and Beyond
arXiv:1904.09237
Abstract
Several recently proposed stochastic optimization methods that have been successfully used in training deep networks such as RMSProp, Adam, Adadelta, Nadam are based on using gradient updates scaled by square roots of exponential moving averages of squared past gradients. In many applications, e.g. learning with large output spaces, it has been empirically observed that these algorithms fail to converge to an optimal solution (or a critical point in nonconvex settings). We show that one cause for such failures is the exponential moving average used in the algorithms. We provide an explicit example of a simple convex optimization setting where Adam does not converge to the optimal solution, and describe the precise problems with the previous analysis of Adam algorithm. Our analysis suggests that the convergence issues can be fixed by endowing such algorithms with `long-term memory' of past gradients, and propose new variants of the Adam algorithm which not only fix the convergence issues but often also lead to improved empirical performance.
Appeared in ICLR 2018
Cited by in corpus (92)
- Improving Generalization Performance by Switching from Adam to SGD
- A Comparison of Optimization Algorithms for Deep Learning
- Fake News Detection on Social Media using Geometric Deep Learning
- How Does Learning Rate Decay Help Modern Neural Networks?
- Optimization for deep learning: theory and algorithms
- Heterogeneous Domain Generalization via Domain Mixup
- Federated Learning with Matched Averaging
- Torchreid: A Library for Deep Learning Person Re-Identification in Pytorch
- Feature-Critic Networks for Heterogeneous Domain Generalization
- Fast simulation of muons produced at the SHiP experiment using Generative Adversarial Networks
- A Selective Overview of Deep Learning
- Equivariant Transformer Networks
- Distributed Training with Heterogeneous Data: Bridging Median- and Mean-Based Algorithms
- Neural-Network Assisted Study of Nitrogen Atom Dynamics on Amorphous Solid Water. I. Adsorption & Desorption
- QUOTIENT: Two-Party Secure Neural Network Training and Prediction
- An Adaptive and Momental Bound Method for Stochastic Learning
- Deep Neural Networks for Marine Debris Detection in Sonar Images
- Riemannian adaptive stochastic gradient algorithms on matrix manifolds
- SGD Converges to Global Minimum in Deep Learning via Star-convex Path
- XNAS: Neural Architecture Search with Expert Advice
- Polysemous Visual-Semantic Embedding for Cross-Modal Retrieval
- Towards Unified INT8 Training for Convolutional Neural Network
- From FiLM to Video: Multi-turn Question Answering with Multi-modal Context
- Does Adam optimizer keep close to the optimal point?
- SRNet: Improving Generalization in 3D Human Pose Estimation with a Split-and-Recombine Approach
- Towards Robust, Locally Linear Deep Networks
- Competence-based Curriculum Learning for Neural Machine Translation
- Temporal-Clustering Invariance in Irregular Healthcare Time Series
- Heterogeneous Molecular Graph Neural Networks for Predicting Molecule Properties
- A High Probability Analysis of Adaptive SGD with Momentum
- WeMix: How to Better Utilize Data Augmentation
- RawNet: Advanced end-to-end deep neural network using raw waveforms for text-independent speaker verification
- Replay attack detection with complementary high-resolution information using end-to-end DNN for the ASVspoof 2019 Challenge
- FlowSeq: Non-Autoregressive Conditional Sequence Generation with Generative Flow
- Anomaly Detection for Industrial Control Systems Using Sequence-to-Sequence Neural Networks
- Scalable Psychological Momentum Forecasting in Esports
- Explicitly Conditioned Melody Generation: A Case Study with Interdependent RNNs
- LiftFormer: 3D Human Pose Estimation using attention models
- Towards Efficient and Unbiased Implementation of Lipschitz Continuity in GANs
- Neural Density Estimation and Likelihood-free Inference
- Disentangling Adaptive Gradient Methods from Learning Rates
- Autoencoder-Based Incremental Class Learning without Retraining on Old Data
- TAdam: A Robust Stochastic Gradient Optimizer
- A General Family of Stochastic Proximal Gradient Methods for Deep Learning
- Stochastic Gradient Descent with Polyak's Learning Rate
- DeepOBS: A Deep Learning Optimizer Benchmark Suite
- Robust and efficient algorithms for high-dimensional black-box quantum optimization
- Adaptive Step Sizes in Variance Reduction via Regularization
- Gradient descent with momentum --- to accelerate or to super-accelerate?
- Conjugate-gradient-based Adam for stochastic optimization and its application to deep learning
- Adaptivity and Optimality: A Universal Algorithm for Online Convex Optimization
- Adaptive Serverless Learning
- Deep Discriminative Representation Learning with Attention Map for Scene Classification
- Mapping in a cycle: Sinkhorn regularized unsupervised learning for point cloud shapes
- Importance Weighted Hierarchical Variational Inference
- Improving Human Text Comprehension through Semi-Markov CRF-based Neural Section Title Generation
- Unsupervised Disentanglement GAN for Domain Adaptive Person Re-Identification
- Exploiting the Full Capacity of Deep Neural Networks while Avoiding Overfitting by Targeted Sparsity Regularization
- Scalable and Accurate Dialogue State Tracking via Hierarchical Sequence Generation
- Adaptively Preconditioned Stochastic Gradient Langevin Dynamics
- A Fully Stochastic Second-Order Trust Region Method
- An Orthogonal-SGD based Learning Approach for MIMO Detection under Multiple Channel Models
- A study on the role of subsidiary information in replay attack spoofing detection
- Unsupervised Learning of Camera Pose with Compositional Re-estimation
- Time-Delay Momentum: A Regularization Perspective on the Convergence and Generalization of Stochastic Momentum for Deep Learning
- Approximated Orthonormal Normalisation in Training Neural Networks
- SAdam: A Variant of Adam for Strongly Convex Functions
- Stochastic Gradient Methods with Block Diagonal Matrix Adaptation
- A Survey on Large-scale Machine Learning
- PoseConvGRU: A Monocular Approach for Visual Ego-motion Estimation by Learning
- CProp: Adaptive Learning Rate Scaling from Past Gradient Conformity
- Power Gradient Descent
- Universal Dependency Parsing from Scratch
- SPI-Optimizer: an integral-Separated PI Controller for Stochastic Optimization
- Multi-vision Attention Networks for On-line Red Jujube Grading
- Discrete Point Flow Networks for Efficient Point Cloud Generation
- Improving Visual Recognition using Ambient Sound for Supervision
- QuantNet: Learning to Quantize by Learning within Fully Differentiable Framework
- Semi-Implicit Back Propagation
- Rapidly Adapting Moment Estimation
- Variance Regularization for Accelerating Stochastic Optimization
- Batch Uniformization for Minimizing Maximum Anomaly Score of DNN-based Anomaly Detection in Sounds
- Convergence of a Relaxed Variable Splitting Coarse Gradient Descent Method for Learning Sparse Weight Binarized Activation Neural Networks
- Hierarchical BiGraph Neural Network as Recommendation Systems
- A Search-based Neural Model for Biomedical Nested and Overlapping Event Detection
- Learning Only from Relevant Keywords and Unlabeled Documents
- Discovering Invariances in Healthcare Neural Networks
- PDE-Inspired Algorithms for Semi-Supervised Learning on Point Clouds
- Adam Induces Implicit Weight Sparsity in Rectifier Neural Networks
- A straightforward line search approach on the expected empirical loss for stochastic deep learning problems
- BAMSProd: A Step towards Generalizing the Adaptive Optimization Methods to Deep Binary Model
- Do place cells dream of conditional probabilities? Learning Neural Nyström representations