Learning to learn by gradient descent by gradient descent
arXiv:1606.04474
Abstract
The move from hand-designed features to learned features in machine learning has been wildly successful. In spite of this, optimization algorithms are still designed by hand. In this paper we show how the design of an optimization algorithm can be cast as a learning problem, allowing the algorithm to learn to exploit structure in the problems of interest in an automatic way. Our learned algorithms, implemented by LSTMs, outperform generic, hand-designed competitors on the tasks for which they are trained, and also generalize well to new tasks with similar structure. We demonstrate this on a number of tasks, including simple convex problems, training neural networks, and styling images with neural art.
References in corpus (1)
Cited by in corpus (28)
- Neural Architecture Search with Reinforcement Learning
- Deep neural networks for the evaluation and design of photonic devices
- Learning to reinforcement learn
- Evolved Policy Gradients
- Automated Reinforcement Learning (AutoRL): A Survey and Open Problems
- Learning to Optimize Variational Quantum Circuits to Solve Combinatorial Problems
- Learning to Draw Samples: With Application to Amortized MLE for Generative Adversarial Learning
- Open-world Learning and Application to Product Classification
- Learning to Optimize Neural Nets
- Learning What Data to Learn
- Goal-Aware Neural SAT Solver
- Learning Deep Energy Models: Contrastive Divergence vs. Amortized MLE
- Stateless Neural Meta-Learning using Second-Order Gradients
- A New Backpropagation Algorithm without Gradient Descent
- Optimising Optimisers with Push GP
- CosmicRIM : Reconstructing Early Universe by Combining Differentiable Simulations with Recurrent Inference Machines
- Training Auto-encoder-based Optimizers for Terahertz Image Reconstruction
- Meta-strategy for Learning Tuning Parameters with Guarantees
- On architectural choices in deep learning: From network structure to gradient convergence and parameter estimation
- Genetic algorithms with DNN-based trainable crossover as an example of partial specialization of general search
- Gradient-based algorithms for multi-objective bi-level optimization
- Meta Learning to Rank for Sparsely Supervised Queries
- Learning the Step-size Policy for the Limited-Memory Broyden-Fletcher-Goldfarb-Shanno Algorithm
- Improving traffic sign recognition by active search
- On the Convergence of Prior-Guided Zeroth-Order Optimization Algorithms
- Meta-Learning a Dynamical Language Model
- AM-Net: Adaptively Aligned Multi-Scale Moment for Few-Shot Action Recognition
- Moco: A Learnable Meta Optimizer for Combinatorial Optimization