The Marginal Value of Adaptive Gradient Methods in Machine Learning
arXiv:1705.08292
Abstract
Adaptive optimization methods, which perform local optimization with a metric constructed from the history of iterates, are becoming increasingly popular for training deep neural networks. Examples include AdaGrad, RMSProp, and Adam. We show that for simple overparameterized problems, adaptive methods often find drastically different solutions than gradient descent (GD) or stochastic gradient descent (SGD). We construct an illustrative binary classification problem where the data is linearly separable, GD and SGD achieve zero test error, and AdaGrad, Adam, and RMSProp attain test errors arbitrarily close to half. We additionally study the empirical generalization capability of adaptive methods on several state-of-the-art deep learning models. We observe that the solutions found by adaptive methods generalize worse (often significantly worse) than SGD, even when these solutions have better training performance. These results suggest that practitioners should reconsider the use of adaptive methods to train neural networks.
References in corpus (5)
- Path-SGD: Path-Normalized Optimization in Deep Neural Networks
- Non-convex learning via Stochastic Gradient Langevin Dynamics: a nonasymptotic analysis
- In Search of the Real Inductive Bias: On the Role of Implicit Regularization in Deep Learning
- Diving into the shallows: a computational perspective on large-scale shallow learning
- Span-Based Constituency Parsing with a Structure-Label System and Provably Optimal Dynamic Oracles
Cited by in corpus (91)
- Don't Decay the Learning Rate, Increase the Batch Size
- Regularizing and Optimizing LSTM Language Models
- Invariant Risk Minimization
- Automated Pavement Crack Segmentation Using U-Net-based Convolutional Neural Network
- Regularization for Deep Learning: A Taxonomy
- Ensemble Kalman Inversion: A Derivative-Free Technique For Machine Learning Tasks
- Optimize TSK Fuzzy Systems for Classification Problems: Mini-Batch Gradient Descent with Uniform Regularization and Batch Normalization
- Review: Deep Learning in Electron Microscopy
- Deep neural networks algorithms for stochastic control problems on finite horizon: convergence analysis
- Understanding the Role of Training Regimes in Continual Learning
- Non-Vacuous Generalization Bounds at the ImageNet Scale: A PAC-Bayesian Compression Approach
- Closing the Generalization Gap of Adaptive Gradient Methods in Training Deep Neural Networks
- Stochastic Gradient Methods with Layer-wise Adaptive Moments for Training of Deep Networks
- Convergence guarantees for RMSProp and ADAM in non-convex optimization and an empirical comparison to Nesterov acceleration
- On the Convergence of Adaptive Gradient Methods for Nonconvex Optimization
- A Closer Look at Deep Learning Heuristics: Learning rate restarts, Warmup and Distillation
- Large Kernel Distillation Network for Efficient Single Image Super-Resolution
- Nostalgic Adam: Weighting more of the past gradients when designing the adaptive learning rate
- Fast and Scalable Bayesian Deep Learning by Weight-Perturbation in Adam
- Towards Theoretically Understanding Why SGD Generalizes Better Than ADAM in Deep Learning
- Reducing Noise in GAN Training with Variance Reduced Extragradient
- WNGrad: Learn the Learning Rate in Gradient Descent
- Hermite-Gaussian Mode Detection via Convolution Neural Networks
- Ten Quick Tips for Deep Learning in Biology
- On the Geometry of Adversarial Examples
- Hamiltonian Descent Methods
- Transfer learning-based method for automated ewaste recycling in smart cities
- Deep Gamblers: Learning to Abstain with Portfolio Theory
- An Adaptive and Momental Bound Method for Stochastic Learning
- Rolling the dice for better deep learning performance: A study of randomness techniques in deep neural networks
- On the Relationship Between Short-Time Objective Intelligibility and Short-Time Spectral-Amplitude Mean-Square Error for Speech Enhancement
- Autoencoding with a Classifier System
- Lexicographic and Depth-Sensitive Margins in Homogeneous and Non-Homogeneous Deep Models
- Riemannian adaptive stochastic gradient algorithms on matrix manifolds
- Skin Cancer Classification using Inception Network and Transfer Learning
- Normalized Direction-preserving Adam
- Near-optimal control of dynamical systems with neural ordinary differential equations
- Algorithmic Regularization in Over-parameterized Matrix Sensing and Neural Networks with Quadratic Activations
- Predicting Adolescent Suicide Attempts with Neural Networks
- On the Convergence of AdaBound and its Connection to SGD
- On the distance between two neural networks and the stability of learning
- Block-Normalized Gradient Method: An Empirical Study for Training Deep Neural Network
- Are we Forgetting about Compositional Optimisers in Bayesian Optimisation?
- Lifted Neural Networks
- Convergence Analysis of Optimization Algorithms
- Reducing the variance in online optimization by transporting past gradients
- Fix your classifier: the marginal value of training the last weight layer
- Visualizing high-dimensional loss landscapes with Hessian directions
- An Empirical Study of Large-Batch Stochastic Gradient Descent with Structured Covariance Noise
- Two layer Ensemble of Deep Learning Models for Medical Image Segmentation
- Second-order step-size tuning of SGD for non-convex optimization
- AdaS: Adaptive Scheduling of Stochastic Gradients
- An Adaptive Gradient Method with Energy and Momentum
- Empirical study towards understanding line search approximations for training neural networks
- Interpreting Deep Learning: The Machine Learning Rorschach Test?
- Escaping Saddle Points Faster with Stochastic Momentum
- The Bayesian Learning Rule
- Using Mode Connectivity for Loss Landscape Analysis
- Weakly Supervised Estimation of Shadow Confidence Maps in Fetal Ultrasound Imaging
- An Adaptive Memory Multi-Batch L-BFGS Algorithm for Neural Network Training
- Towards Model Agnostic Federated Learning Using Knowledge Distillation
- Reinforced stochastic gradient descent for deep neural network learning
- Revisiting Recursive Least Squares for Training Deep Neural Networks
- Automatic quantification of abdominal subcutaneous and visceral adipose tissue in children, through MRI study, using total intensity maps and Convolutional Neural Networks
- Gravity Optimizer: a Kinematic Approach on Optimization in Deep Learning
- How Much Restricted Isometry is Needed In Nonconvex Matrix Recovery?
- Learning scale-variant features for robust iris authentication with deep learning based ensemble framework
- Variance Reduction on General Adaptive Stochastic Mirror Descent
- Training Feedforward Neural Networks with Standard Logistic Activations is Feasible
- Stochastic natural gradient descent draws posterior samples in function space
- Structured Dropout Variational Inference for Bayesian Neural Networks
- A Self-Attention-Driven Deep Denoiser Model for Real Time Lung Sound Denoising in Noisy Environments
- GABO: Graph Augmentations with Bi-level Optimization
- Logit Attenuating Weight Normalization
- Painless step size adaptation for SGD
- Reappraising Domain Generalization in Neural Networks
- Towards Robust and Automatic Hyper-Parameter Tunning
- Vanishing Curvature and the Power of Adaptive Methods in Randomly Initialized Deep Networks
- Scaling transition from momentum stochastic gradient descent to plain stochastic gradient descent
- Nonlinear Matrix Approximation with Radial Basis Function Components
- Neograd: Near-Ideal Gradient Descent
- Backtracking gradient descent method for general functions, with applications to Deep Learning
- Improving Hyperparameter Optimization by Planning Ahead
- On Generalization of Adaptive Methods for Over-parameterized Linear Regression
- Print Error Detection using Convolutional Neural Networks
- Tangent Space Sensitivity and Distribution of Linear Regions in ReLU Networks
- ACMo: Angle-Calibrated Moment Methods for Stochastic Optimization
- Why to "grow" and "harvest" deep learning models?
- Adaptive versus Standard Descent Methods and Robustness Against Adversarial Examples
- Adam Induces Implicit Weight Sparsity in Rectifier Neural Networks
- Neural Sequence Model Training via -divergence Minimization