On the Convergence of Adam and Beyond
arXiv:1904.09237
Abstract
Several recently proposed stochastic optimization methods that have been successfully used in training deep networks such as RMSProp, Adam, Adadelta, Nadam are based on using gradient updates scaled by square roots of exponential moving averages of squared past gradients. In many applications, e.g. learning with large output spaces, it has been empirically observed that these algorithms fail to converge to an optimal solution (or a critical point in nonconvex settings). We show that one cause for such failures is the exponential moving average used in the algorithms. We provide an explicit example of a simple convex optimization setting where Adam does not converge to the optimal solution, and describe the precise problems with the previous analysis of Adam algorithm. Our analysis suggests that the convergence issues can be fixed by endowing such algorithms with `long-term memory' of past gradients, and propose new variants of the Adam algorithm which not only fix the convergence issues but often also lead to improved empirical performance.
Appeared in ICLR 2018
Cited by in corpus (200)
- Improving Generalization Performance by Switching from Adam to SGD
- Robust Motion In-betweening
- A Comparison of Optimization Algorithms for Deep Learning
- Fake News Detection on Social Media using Geometric Deep Learning
- How Does Learning Rate Decay Help Modern Neural Networks?
- Optimization for deep learning: theory and algorithms
- DGCN: Diversified Recommendation with Graph Convolutional Networks
- Heterogeneous Domain Generalization via Domain Mixup
- Gaussian Moments as Physically Inspired Molecular Descriptors for Accurate and Scalable Machine Learning Potentials
- Federated Learning with Matched Averaging
- A Comparison of Various Classical Optimizers for a Variational Quantum Linear Solver
- Torchreid: A Library for Deep Learning Person Re-Identification in Pytorch
- Feature-Critic Networks for Heterogeneous Domain Generalization
- Large Kernel Distillation Network for Efficient Single Image Super-Resolution
- Fast simulation of muons produced at the SHiP experiment using Generative Adversarial Networks
- Adam revisited: a weighted past gradients perspective
- Light Field Image Quality Assessment With Auxiliary Learning Based on Depthwise and Anglewise Separable Convolutions
- A Selective Overview of Deep Learning
- Distributed Training with Heterogeneous Data: Bridging Median- and Mean-Based Algorithms
- Equivariant Transformer Networks
- Neural-Network Assisted Study of Nitrogen Atom Dynamics on Amorphous Solid Water. I. Adsorption & Desorption
- QUOTIENT: Two-Party Secure Neural Network Training and Prediction
- An Adaptive and Momental Bound Method for Stochastic Learning
- Towards Understanding Label Smoothing
- Deep Neural Networks for Marine Debris Detection in Sonar Images
- Group Convolutional Neural Networks Improve Quantum State Accuracy
- A deep learning driven pseudospectral PCE based FFT homogenization algorithm for complex microstructures
- SE(3)-equivariant prediction of molecular wavefunctions and electronic densities
- AdaX: Adaptive Gradient Descent with Exponential Long Term Memory
- Riemannian adaptive stochastic gradient algorithms on matrix manifolds
- XNAS: Neural Architecture Search with Expert Advice
- Polysemous Visual-Semantic Embedding for Cross-Modal Retrieval
- SGD Converges to Global Minimum in Deep Learning via Star-convex Path
- Local Adaptivity in Federated Learning: Convergence and Consistency
- STEM: A Stochastic Two-Sided Momentum Algorithm Achieving Near-Optimal Sample and Communication Complexities for Federated Learning
- Towards Unified INT8 Training for Convolutional Neural Network
- From FiLM to Video: Multi-turn Question Answering with Multi-modal Context
- Does Adam optimizer keep close to the optimal point?
- Mixing ADAM and SGD: a Combined Optimization Method
- SRNet: Improving Generalization in 3D Human Pose Estimation with a Split-and-Recombine Approach
- Towards Robust, Locally Linear Deep Networks
- Classification of Upper Arm Movements from EEG signals using Machine Learning with ICA Analysis
- Competence-based Curriculum Learning for Neural Machine Translation
- Learning to Optimize in Model Predictive Control
- Detection of Illicit Drug Trafficking Events on Instagram: A Deep Multimodal Multilabel Learning Approach
- Estimation of discrete choice models with hybrid stochastic adaptive batch size algorithms
- Temporal-Clustering Invariance in Irregular Healthcare Time Series
- A High Probability Analysis of Adaptive SGD with Momentum
- Neural Pitch-Shifting and Time-Stretching with Controllable LPCNet
- Heterogeneous Molecular Graph Neural Networks for Predicting Molecule Properties
- WeMix: How to Better Utilize Data Augmentation
- Replay attack detection with complementary high-resolution information using end-to-end DNN for the ASVspoof 2019 Challenge
- Adam: A Stochastic Method with Adaptive Variance Reduction
- RawNet: Advanced end-to-end deep neural network using raw waveforms for text-independent speaker verification
- Hyperbolic Deep Neural Networks: A Survey
- FlowSeq: Non-Autoregressive Conditional Sequence Generation with Generative Flow
- Anomaly Detection for Industrial Control Systems Using Sequence-to-Sequence Neural Networks
- LiftFormer: 3D Human Pose Estimation using attention models
- Explicitly Conditioned Melody Generation: A Case Study with Interdependent RNNs
- Accelerated Large Batch Optimization of BERT Pretraining in 54 minutes
- Non-Convergence and Limit Cycles in the Adam optimizer
- Scalable Psychological Momentum Forecasting in Esports
- AdaS: Adaptive Scheduling of Stochastic Gradients
- Large-Scale Deep Learning Optimizations: A Comprehensive Survey
- Towards Efficient and Unbiased Implementation of Lipschitz Continuity in GANs
- Stochastic Gradient Descent with Nonlinear Conjugate Gradient-Style Adaptive Momentum
- An Adaptive Gradient Method with Energy and Momentum
- TAdam: A Robust Stochastic Gradient Optimizer
- Adaptive First-and Zeroth-order Methods for Weakly Convex Stochastic Optimization Problems
- SceneGraphFusion: Incremental 3D Scene Graph Prediction from RGB-D Sequences
- Private Stochastic Non-Convex Optimization: Adaptive Algorithms and Tighter Generalization Bounds
- Disentangling Adaptive Gradient Methods from Learning Rates
- Autoencoder-Based Incremental Class Learning without Retraining on Old Data
- Neural Density Estimation and Likelihood-free Inference
- Stochastic Gradient Descent with Polyak's Learning Rate
- Improved RawNet with Feature Map Scaling for Text-independent Speaker Verification using Raw Waveforms
- ARC: A Vision-based Automatic Retail Checkout System
- A Decentralized Adaptive Momentum Method for Solving a Class of Min-Max Optimization Problems
- A General Family of Stochastic Proximal Gradient Methods for Deep Learning
- Escaping Saddle Points Faster with Stochastic Momentum
- The Role of Momentum Parameters in the Optimal Convergence of Adaptive Polyak's Heavy-ball Methods
- Understanding the Generalization of Adam in Learning Neural Networks with Proper Regularization
- DeepOBS: A Deep Learning Optimizer Benchmark Suite
- Robust and efficient algorithms for high-dimensional black-box quantum optimization
- On the Effectiveness of Regularization Against Membership Inference Attacks
- Phase-aware Single-stage Speech Denoising and Dereverberation with U-Net
- Adaptive Step Sizes in Variance Reduction via Regularization
- AdaSGD: Bridging the gap between SGD and Adam
- Exploiting Adam-like Optimization Algorithms to Improve the Performance of Convolutional Neural Networks
- On the Convergence of Decentralized Adaptive Gradient Methods
- Conjugate-gradient-based Adam for stochastic optimization and its application to deep learning
- Adaptivity and Optimality: A Universal Algorithm for Online Convex Optimization
- Gradient descent with momentum --- to accelerate or to super-accelerate?
- MixML: A Unified Analysis of Weakly Consistent Parallel Learning
- Stability of SGD: Tightness Analysis and Improved Bounds
- On the Distributional Properties of Adaptive Gradients
- An Element-Wise Weights Aggregation Method for Federated Learning
- Learning from Imperfect Annotations
- The Nonlinearity Coefficient - A Practical Guide to Neural Architecture Design
- Exploring Severe Occlusion: Multi-Person 3D Pose Estimation with Gated Convolution
- Sinkhorn Natural Gradient for Generative Models
- Deep Discriminative Representation Learning with Attention Map for Scene Classification
- Early Detection of Fake News by Utilizing the Credibility of News, Publishers, and Users Based on Weakly Supervised Learning
- Convergence of adaptive algorithms for weakly convex constrained optimization
- Adaptive Serverless Learning
- Scalable and Accurate Dialogue State Tracking via Hierarchical Sequence Generation
- A Fully Stochastic Second-Order Trust Region Method
- Unsupervised Disentanglement GAN for Domain Adaptive Person Re-Identification
- Mapping in a cycle: Sinkhorn regularized unsupervised learning for point cloud shapes
- Exploiting the Full Capacity of Deep Neural Networks while Avoiding Overfitting by Targeted Sparsity Regularization
- Multi-Channel and Multi-Microphone Acoustic Echo Cancellation Using A Deep Learning Based Approach
- Rethinking movie genre classification with fine-grained semantic clustering
- Representing Partial Programs with Blended Abstract Semantics
- Doubly Adaptive Scaled Algorithm for Machine Learning Using Second-Order Information
- Improving Human Text Comprehension through Semi-Markov CRF-based Neural Section Title Generation
- Importance Weighted Hierarchical Variational Inference
- High-probability Bounds for Non-Convex Stochastic Optimization with Heavy Tails
- Adaptively Preconditioned Stochastic Gradient Langevin Dynamics
- The Development of Spatial Attention U-Net for The Recovery of Ionospheric Measurements and The Extraction of Ionospheric Parameters
- -continuous Spline Approximation with TensorFlow Gradient Descent Optimizers
- SAdam: A Variant of Adam for Strongly Convex Functions
- Optimizer Fusion: Efficient Training with Better Locality and Parallelism
- Approximated Orthonormal Normalisation in Training Neural Networks
- Unsupervised Learning of Camera Pose with Compositional Re-estimation
- PoseConvGRU: A Monocular Approach for Visual Ego-motion Estimation by Learning
- A Survey on Large-scale Machine Learning
- On the One-sided Convergence of Adam-type Algorithms in Non-convex Non-concave Min-max Optimization
- Incorporating the Barzilai-Borwein Adaptive Step Size into Sugradient Methods for Deep Network Training
- A study on the role of subsidiary information in replay attack spoofing detection
- Stochastic Gradient Methods with Block Diagonal Matrix Adaptation
- Adaptive Differentially Private Empirical Risk Minimization
- Time-Delay Momentum: A Regularization Perspective on the Convergence and Generalization of Stochastic Momentum for Deep Learning
- Tensor-Train Networks for Learning Predictive Modeling of Multidimensional Data
- An Orthogonal-SGD based Learning Approach for MIMO Detection under Multiple Channel Models
- Gravity Optimizer: a Kinematic Approach on Optimization in Deep Learning
- Microscopy Image Restoration using Deep Learning on W2S
- Applying Cyclical Learning Rate to Neural Machine Translation
- Adaptive Gradient Methods Can Be Provably Faster than SGD after Finite Epochs
- Training Aware Sigmoidal Optimizer
- Segment Aggregation for short utterances speaker verification using raw waveforms
- Generalizing Adversarial Examples by AdaBelief Optimizer
- -Neighbor Based Curriculum Sampling for Sequence Prediction
- Loss Landscape Dependent Self-Adjusting Learning Rates in Decentralized Stochastic Gradient Descent
- SPI-Optimizer: an integral-Separated PI Controller for Stochastic Optimization
- CADA: Communication-Adaptive Distributed Adam
- Solving Differential Equations via Continuous-Variable Quantum Computers
- Universal Dependency Parsing from Scratch
- Competition analysis on the over-the-counter credit default swap market
- Discrete Point Flow Networks for Efficient Point Cloud Generation
- PCLs: Geometry-aware Neural Reconstruction of 3D Pose with Perspective Crop Layers
- Anatomical and Diagnostic Bayesian Segmentation in Prostate MRI Should Different Clinical Objectives Mandate Different Loss Functions?
- Multi-vision Attention Networks for On-line Red Jujube Grading
- L2M: Practical posterior Laplace approximation with optimization-driven second moment estimation
- AdaL: Adaptive Gradient Transformation Contributes to Convergences and Generalizations
- CProp: Adaptive Learning Rate Scaling from Past Gradient Conformity
- Data Augmentation for Text Generation Without Any Augmented Data
- Power Gradient Descent
- Compressed Communication for Distributed Training: Adaptive Methods and System
- Extraction of Hierarchical Functional Connectivity Components in human brain using Adversarial Learning
- Spatio-Temporal Neural Network for Fitting and Forecasting COVID-19
- Universal Representation Learning of Knowledge Bases by Jointly Embedding Instances and Ontological Concepts
- Distributed Stochastic Non-Convex Optimization: Momentum-Based Variance Reduction
- Variance Regularization for Accelerating Stochastic Optimization
- The Number of Steps Needed for Nonconvex Optimization of a Deep Learning Optimizer is a Rational Function of Batch Size
- Accelerating Distributed SGD for Linear Regression using Iterative Pre-Conditioning
- Identifying Illicit Drug Dealers on Instagram with Large-scale Multimodal Data Fusion
- Efficient Semi-Implicit Variational Inference
- A Search-based Neural Model for Biomedical Nested and Overlapping Event Detection
- Adaptive Learning Rate and Momentum for Training Deep Neural Networks
- A New Adaptive Gradient Method with Gradient Decomposition
- Learning Only from Relevant Keywords and Unlabeled Documents
- Local Convergence of Adaptive Gradient Descent Optimizers
- Learning Metrics from Mean Teacher: A Supervised Learning Method for Improving the Generalization of Speaker Verification System
- QuantNet: Learning to Quantize by Learning within Fully Differentiable Framework
- Understanding Modern Techniques in Optimization: Frank-Wolfe, Nesterov's Momentum, and Polyak's Momentum
- Automatic Live Music Song Identification Using Multi-level Deep Sequence Similarity Learning
- Improving Visual Recognition using Ambient Sound for Supervision
- Scaling transition from momentum stochastic gradient descent to plain stochastic gradient descent
- AsymptoticNG: A regularized natural gradient optimization algorithm with look-ahead strategy
- A Comprehensive Study on Optimization Strategies for Gradient Descent In Deep Learning
- PDE-Inspired Algorithms for Semi-Supervised Learning on Point Clouds
- Do place cells dream of conditional probabilities? Learning Neural Nyström representations
- Neural representation and generation for RNA secondary structures
- Adam Induces Implicit Weight Sparsity in Rectifier Neural Networks
- 2nd-order Updates with 1st-order Complexity
- BAMSProd: A Step towards Generalizing the Adaptive Optimization Methods to Deep Binary Model
- A straightforward line search approach on the expected empirical loss for stochastic deep learning problems
- Batch Uniformization for Minimizing Maximum Anomaly Score of DNN-based Anomaly Detection in Sounds
- Rapidly Adapting Moment Estimation
- Hierarchical BiGraph Neural Network as Recommendation Systems
- Discovering Invariances in Healthcare Neural Networks
- :An Unbiased Stratified Statistic and a Fast Gradient Optimization Algorithm Based on It
- A Study of Policy Gradient on a Class of Exactly Solvable Models
- Semi-Implicit Back Propagation
- Active Learning for Argument Mining: A Practical Approach
- Convergence of a Relaxed Variable Splitting Coarse Gradient Descent Method for Learning Sparse Weight Binarized Activation Neural Networks
- Learning Constraints and Descriptive Segmentation for Subevent Detection
- On Generalization of Adaptive Methods for Over-parameterized Linear Regression
- Toward Communication Efficient Adaptive Gradient Method
- Tom: Leveraging trend of the observed gradients for faster convergence