Descending through a Crowded Valley - Benchmarking Deep Learning Optimizers
arXiv:2007.01547
Abstract
Choosing the optimizer is considered to be among the most crucial design decisions in deep learning, and it is not an easy one. The growing literature now lists hundreds of optimization methods. In the absence of clear theoretical guidance and conclusive empirical evidence, the decision is often made based on anecdotes. In this work, we aim to replace these anecdotes, if not with a conclusive ranking, then at least with evidence-backed heuristics. To do so, we perform an extensive, standardized benchmark of fifteen particularly popular deep learning optimizers while giving a concise overview of the wide range of possible choices. Analyzing more than individual runs, we contribute the following three points: (i) Optimizer performance varies greatly across tasks. (ii) We observe that evaluating multiple optimizers with default parameters works approximately as well as tuning the hyperparameters of a single, fixed optimizer. (iii) While we cannot discern an optimization method clearly dominating across all tested tasks, we identify a significantly reduced subset of specific optimizers and parameter choices that generally lead to competitive results in our experiments: Adam remains a strong contender, with newer methods failing to significantly and consistently outperform it. Our open-sourced results are available as challenging and well-tuned baselines for more meaningful evaluations of novel optimization methods without requiring any further computational efforts.
Raw results: https://github.com/SirRob1997/Crowded-Valley---Results
References in corpus (46)
- ADADELTA: An Adaptive Learning Rate Method
- On the Convergence of Adam and Beyond
- Large Batch Training of Convolutional Networks
- Improving Generalization Performance by Switching from Adam to SGD
- No More Pesky Learning Rates
- Neural Optimizer Search with Reinforcement Learning
- Adaptive Gradient Methods with Dynamic Bound of Learning Rate
- AdaComp : Adaptive Residual Gradient Compression for Data-Parallel Distributed Training
- Adaptivity without Compromise: A Momentumized, Adaptive, Dual Averaged Gradient Method for Stochastic Optimization
- Practical Gauss-Newton Optimisation for Deep Learning
- An Adaptive and Momental Bound Method for Stochastic Learning
- AdaX: Adaptive Gradient Descent with Exponential Long Term Memory
- PAGE: A Simple and Optimal Probabilistic Gradient Estimator for Nonconvex Optimization
- AngularGrad: A New Optimization Technique for Angular Convergence of Convolutional Neural Networks
- AdaScale SGD: A User-Friendly Algorithm for Distributed Training
- ASAM: Adaptive Sharpness-Aware Minimization for Scale-Invariant Learning of Deep Neural Networks
- Does Adam optimizer keep close to the optimal point?
- Mixing ADAM and SGD: a Combined Optimization Method
- Adaptive learning rates and parallelization for stochastic, sparse, non-smooth gradients
- CoolMomentum: A Method for Stochastic Optimization by Langevin Dynamics with Simulated Annealing
- Second-order step-size tuning of SGD for non-convex optimization
- AdaS: Adaptive Scheduling of Stochastic Gradients
- Stochastic Gradient Descent with Nonlinear Conjugate Gradient-Style Adaptive Momentum
- TAdam: A Robust Stochastic Gradient Optimizer
- Momentum-based variance-reduced proximal stochastic gradient method for composite nonconvex stochastic optimization
- AutoLRS: Automatic Learning-Rate Schedule by Bayesian Optimization on the Fly
- S-SGD: Symmetrical Stochastic Gradient Descent with Weight Noise Injection for Reaching Flat Minima
- Learning by Turning: Neural Architecture Aware Optimisation
- AdaSGD: Bridging the gap between SGD and Adam
- A Generalizable Approach to Learning Optimizers
- Momentum with Variance Reduction for Nonconvex Composition Optimization
- Expectigrad: Fast Stochastic Optimization with Robust Convergence Properties
- Stochastic Runge-Kutta methods and adaptive SGD-G2 stochastic gradient descent
- Adaptive Gradient Methods Can Be Provably Faster than SGD after Finite Epochs
- Gravity Optimizer: a Kinematic Approach on Optimization in Deep Learning
- Gravilon: Applications of a New Gradient Descent Method to Machine Learning
- A New Accelerated Stochastic Gradient Method with Momentum
- Stochastic Gradient Methods with Block Diagonal Matrix Adaptation
- ICLR Reproducibility Challenge Report (Padam : Closing The Generalization Gap Of Adaptive Gradient Methods in Training Deep Neural Networks)
- On Higher-order Moments in Adam
- A Probabilistically Motivated Learning Rate Adaptation for Stochastic Optimization
- CProp: Adaptive Learning Rate Scaling from Past Gradient Conformity
- A Variant of Gradient Descent Algorithm Based on Gradient Averaging
- Learning in Deep Neural Networks Using a Biologically Inspired Optimizer
- Second-order Information in First-order Optimization Methods
- GOALS: Gradient-Only Approximations for Line Searches Towards Robust and Consistent Training of Deep Neural Networks
Cited by in corpus (16)
- Review: Deep Learning in Electron Microscopy
- Adaptivity without Compromise: A Momentumized, Adaptive, Dual Averaged Gradient Method for Stochastic Optimization
- Are we Forgetting about Compositional Optimisers in Bayesian Optimisation?
- A Large Batch Optimizer Reality Check: Traditional, Generic Optimizers Suffice Across Batch Sizes
- Stochastic Training is Not Necessary for Generalization
- FCM-RDpA: TSK Fuzzy Regression Model Construction Using Fuzzy C-Means Clustering, Regularization, DropRule, and Powerball AdaBelief
- Ranger21: a synergistic deep learning optimizer
- Deep Neural Network Training with Frank-Wolfe
- Deep learning for the modeling and inverse design of radiative heat transfer
- Meta-Learning Bidirectional Update Rules
- Explainability-aided Domain Generalization for Image Classification
- Length Scale Control in Topology Optimization using Fourier Enhanced Neural Networks
- Training Aware Sigmoidal Optimizer
- Surrogate Model Based Hyperparameter Tuning for Deep Learning with SPOT
- A straightforward line search approach on the expected empirical loss for stochastic deep learning problems
- The Number of Steps Needed for Nonconvex Optimization of a Deep Learning Optimizer is a Rational Function of Batch Size