A Generalizable Approach to Learning Optimizers
arXiv:2106.00958
Abstract
A core issue with learning to optimize neural networks has been the lack of generalization to real world problems. To address this, we describe a system designed from a generalization-first perspective, learning to update optimizer hyperparameters instead of model parameters directly using novel features, actions, and a reward function. This system outperforms Adam at all neural network tasks including on modalities not seen during training. We achieve 2x speedups on ImageNet, and a 2.5x speedup on a language modeling task using over 5 orders of magnitude more compute than the training tasks.
References in corpus (14)
- PyTorch: An Imperative Style, High-Performance Deep Learning Library
- Fashion-MNIST: a Novel Image Dataset for Benchmarking Machine Learning Algorithms
- Language Models are Few-Shot Learners
- Scaling Laws for Neural Language Models
- Large Batch Training of Convolutional Networks
- Pointer Sentinel Mixture Models
- Gradient-based Hyperparameter Optimization through Reversible Learning
- Practical recommendations for gradient-based training of deep architectures
- MLPerf Training Benchmark
- An Empirical Model of Large-Batch Training
- Single Headed Attention RNN: Stop Thinking With Your Head
- Reinforcement Learning for Learning Rate Control
- Tasks, stability, architecture, and compute: Training more effective learned optimizers, and using them to train themselves
- The Two Regimes of Deep Network Training