Learned Optimizers that Scale and Generalize
arXiv:1703.04813
Abstract
Learning to learn has emerged as an important direction for achieving artificial intelligence. Two of the primary barriers to its adoption are an inability to scale to larger problems and a limited ability to generalize to new tasks. We introduce a learned gradient descent optimizer that generalizes well to new tasks, and which has significantly reduced memory and computation overhead. We achieve this by introducing a novel hierarchical RNN architecture, with minimal per-parameter overhead, augmented with additional architectural features that mirror the known structure of optimization tasks. We also develop a meta-training ensemble of small, diverse optimization tasks capturing common properties of loss landscapes. The optimizer learns to outperform RMSProp/ADAM on problems in this corpus. More importantly, it performs comparably or better when applied to small convolutional neural networks, despite seeing no neural networks in its meta-training set. Finally, it generalizes to train Inception V3 and ResNet V2 architectures on the ImageNet dataset for thousands of steps, optimization problems that are of a vastly different scale than those it was trained on. We release an open source implementation of the meta-training algorithm.
Final ICML paper after reviewer suggestions
References in corpus (3)
Cited by in corpus (42)
- Learning to Optimize: A Primer and A Benchmark
- Meta-Learning Update Rules for Unsupervised Representation Learning
- Scalable Neural Architecture Search for 3D Medical Image Segmentation
- Mitigating Metaphors: A Comprehensible Guide to Recent Nature-Inspired Algorithms
- From Learning to Meta-Learning: Reduced Training Overhead and Complexity for Communication Systems
- Understanding Short-Horizon Bias in Stochastic Meta-Optimization
- A Comprehensive Overview and Survey of Recent Advances in Meta-Learning
- Meta Continual Learning
- Deep Frank-Wolfe For Neural Network Optimization
- Learned Robust PCA: A Scalable Deep Unfolding Approach for High-Dimensional Outlier Detection
- BADGER: Learning to (Learn [Learning Algorithms] through Multi-Agent Communication)
- How to Train Your MAML to Excel in Few-Shot Classification
- How to decay your learning rate
- Tasks, stability, architecture, and compute: Training more effective learned optimizers, and using them to train themselves
- Using a thousand optimization tasks to learn hyperparameter search strategies
- Learning to Optimize in Swarms
- Using learned optimizers to make models robust to input noise
- MRL: Mind-aware Multi-agent Management Reinforcement Learning
- Optimising Optimisers with Push GP
- A General Decoupled Learning Framework for Parameterized Image Operators
- Learning to Learn by Zeroth-Order Oracle
- Guarantees for Tuning the Step Size using a Learning-to-Learn Approach
- Training Stronger Baselines for Learning to Optimize
- A Generalizable Approach to Learning Optimizers
- Contextualizing Enhances Gradient Based Meta Learning
- Improved Adversarial Training via Learned Optimizer
- Parameter Prediction for Unseen Deep Architectures
- Meta-Learning Bidirectional Update Rules
- Training Learned Optimizers with Randomly Initialized Learned Optimizers
- Gradients are Not All You Need
- Soft Layer Selection with Meta-Learning for Zero-Shot Cross-Lingual Transfer
- Neural Fixed-Point Acceleration for Convex Optimization
- Model-Agnostic Meta-Attack: Towards Reliable Evaluation of Adversarial Robustness
- Confusable Learning for Large-class Few-Shot Classification
- MetalGAN: a Cluster-based Adaptive Training for Few-Shot Adversarial Colorization
- Accelerating Gradient-based Meta Learner
- ModelPred: A Framework for Predicting Trained Model from Training Data
- Learning to Initialize Gradient Descent Using Gradient Descent
- FISAR: Forward Invariant Safe Reinforcement Learning with a Deep Neural Network-Based Optimize
- Adaptive Hierarchical Hyper-gradient Descent
- Task Attended Meta-Learning for Few-Shot Learning
- Learning to Learn End-to-End Goal-Oriented Dialog From Related Dialog Tasks