Meta-Learning and Universality: Deep Representations and Gradient Descent can Approximate any Learning Algorithm
arXiv:1710.11622
Abstract
Learning to learn is a powerful paradigm for enabling models to learn from data more effectively and efficiently. A popular approach to meta-learning is to train a recurrent model to read in a training dataset as input and output the parameters of a learned model, or output predictions for new test inputs. Alternatively, a more recent approach to meta-learning aims to acquire deep representations that can be effectively fine-tuned, via standard gradient descent, to new tasks. In this paper, we consider the meta-learning problem from the perspective of universality, formalizing the notion of learning algorithm approximation and comparing the expressive power of the aforementioned recurrent models to the more recent approaches that embed gradient descent into the meta-learner. In particular, we seek to answer the following question: does deep representation combined with standard gradient descent have sufficient capacity to approximate any learning algorithm? We find that this is indeed true, and further find, in our experiments, that gradient-based meta-learning consistently leads to learning strategies that generalize more widely compared to those represented by recurrent models.
ICLR 2018
References in corpus (6)
Cited by in corpus (43)
- On First-Order Meta-Learning Algorithms
- Meta-Reinforcement Learning of Structured Exploration Strategies
- Meta-learning with differentiable closed-form solvers
- Evolved Policy Gradients
- Learning to Learn without Forgetting by Maximizing Transfer and Minimizing Interference
- Learning to Adapt in Dynamic, Real-World Environments Through Meta-Reinforcement Learning
- Rapid Learning or Feature Reuse? Towards Understanding the Effectiveness of MAML
- Universal Planning Networks
- A novel meta-learning initialization method for physics-informed neural networks
- Machine Theory of Mind
- Small Sample Learning in Big Data Era
- AI-GAs: AI-generating algorithms, an alternate paradigm for producing general artificial intelligence
- Generative Teaching Networks: Accelerating Neural Architecture Search by Learning to Generate Synthetic Training Data
- Learning from Few Samples: A Survey
- Few-Shot Goal Inference for Visuomotor Learning and Planning
- Improved Contrastive Divergence Training of Energy Based Models
- Adversarial Meta-Learning
- Auto-Meta: Automated Gradient Based Meta Learner Search
- Meta-learning Based Beamforming Design for MISO Downlink
- Probabilistic Model-Agnostic Meta-Learning
- Evolving Reinforcement Learning Algorithms
- Learning to reinforcement learn for Neural Architecture Search
- On the Importance of Attention in Meta-Learning for Few-Shot Text Classification
- Learning Fast Adaptation with Meta Strategy Optimization
- Improving Generalization in Meta-learning via Task Augmentation
- Modeling and Optimization Trade-off in Meta-learning
- Stateless Neural Meta-Learning using Second-Order Gradients
- Is Support Set Diversity Necessary for Meta-Learning?
- ARCADe: A Rapid Continual Anomaly Detector
- Online Structured Meta-learning
- VFunc: a Deep Generative Model for Functions
- Nonstationary Nonparametric Online Learning: Balancing Dynamic Regret and Model Parsimony
- Learning to Learn with Feedback and Local Plasticity
- Graph Few-shot Learning via Knowledge Transfer
- A Model-based Approach for Sample-efficient Multi-task Reinforcement Learning
- Contextualizing Enhances Gradient Based Meta Learning
- MAME : Model-Agnostic Meta-Exploration
- Proximal Mapping for Deep Regularization
- Top-Related Meta-Learning Method for Few-Shot Object Detection
- Is the Meta-Learning Idea Able to Improve the Generalization of Deep Neural Networks on the Standard Supervised Learning?
- Covariate Distribution Aware Meta-learning
- Universality of Gradient Descent Neural Network Training
- Relation-aware Meta-learning for Market Segment Demand Prediction with Limited Records