A Simple Neural Attentive Meta-Learner
arXiv:1707.03141
Abstract
Deep neural networks excel in regimes with large amounts of data, but tend to struggle when data is scarce or when they need to adapt quickly to changes in the task. In response, recent work in meta-learning proposes training a meta-learner on a distribution of similar tasks, in the hopes of generalization to novel but related tasks by learning a high-level strategy that captures the essence of the problem it is asked to solve. However, many recent meta-learning approaches are extensively hand-designed, either using architectures specialized to a particular application, or hard-coding algorithmic components that constrain how the meta-learner solves the task. We propose a class of simple and generic meta-learner architectures that use a novel combination of temporal convolutions and soft attention; the former to aggregate information from past experience and the latter to pinpoint specific pieces of information. In the most extensive set of meta-learning experiments to date, we evaluate the resulting Simple Neural AttentIve Learner (or SNAIL) on several heavily-benchmarked tasks. On all tasks, in both supervised and reinforcement learning, SNAIL attains state-of-the-art performance by significant margins.
iclr 2018 version
Cited by in corpus (255)
- Generalizing from a Few Examples: A Survey on Few-Shot Learning
- Few-Shot Learning with Graph Neural Networks
- Meta-World: A Benchmark and Evaluation for Multi-Task and Meta Reinforcement Learning
- SimpleShot: Revisiting Nearest-Neighbor Classification for Few-Shot Learning
- TADAM: Task dependent adaptive metric for improved few-shot learning
- A Survey of Zero-shot Generalisation in Deep Reinforcement Learning
- Anomaly Detection-Inspired Few-Shot Medical Image Segmentation Through Self-Supervision With Supervoxels
- Bayesian Model-Agnostic Meta-Learning
- A Two-Stage Approach to Few-Shot Learning for Image Recognition
- Automated Machine Learning: State-of-The-Art and Open Challenges
- Set Transformer: A Framework for Attention-based Permutation-Invariant Neural Networks
- Reinforcement Learning for Angle-Only Intercept Guidance of Maneuvering Targets
- Evolved Policy Gradients
- Adaptive Guidance and Integrated Navigation with Reinforcement Meta-Learning
- MetaKG: Meta-learning on Knowledge Graph for Cold-start Recommendation
- Bilevel Programming for Hyperparameter Optimization and Meta-Learning
- Incremental Object Detection via Meta-Learning
- Deep Meta-Learning: Learning to Learn in the Concept Space
- Learning to Adapt in Dynamic, Real-World Environments Through Meta-Reinforcement Learning
- Finding Task-Relevant Features for Few-Shot Learning by Category Traversal
- Seeker based Adaptive Guidance via Reinforcement Meta-Learning Applied to Asteroid Close Proximity Operations
- Meta-Learning and Universality: Deep Representations and Gradient Descent can Approximate any Learning Algorithm
- Rethinking Few-Shot Image Classification: a Good Embedding Is All You Need?
- Efficient Exploration via State Marginal Matching
- Meta-Learning Update Rules for Unsupervised Representation Learning
- Gotta Learn Fast: A New Benchmark for Generalization in RL
- Neuron Linear Transformation: Modeling the Domain Shift for Crowd Counting
- Early Melanoma Diagnosis with Sequential Dermoscopic Images
- Dynamic Few-Shot Visual Learning without Forgetting
- Self-supervised Knowledge Distillation for Few-shot Learning
- Few-Shot Learning with Geometric Constraints
- Meta-Learning with Warped Gradient Descent
- Few-shot Relation Extraction via Bayesian Meta-learning on Relation Graphs
- Bayesian Meta-Learning for the Few-Shot Setting via Deep Kernels
- Improving Generalization in Meta Reinforcement Learning using Learned Objectives
- CrossTransformers: spatially-aware few-shot transfer
- Six Degree-of-Freedom Body-Fixed Hovering over Unmapped Asteroids via LIDAR Altimetry and Reinforcement Meta-Learning
- Meta reinforcement learning as task inference
- Small Sample Learning in Big Data Era
- Unsupervised Meta-Learning for Reinforcement Learning
- GeneraLight: Improving Environment Generalization of Traffic Signal Control via Meta Reinforcement Learning
- Transductive Information Maximization For Few-Shot Learning
- Revisiting Local Descriptor based Image-to-Class Measure for Few-shot Learning
- A Comprehensive Overview and Survey of Recent Advances in Meta-Learning
- Many-Class Few-Shot Learning on Multi-Granularity Class Hierarchy
- MT-Opt: Continuous Multi-Task Robotic Reinforcement Learning at Scale
- Prior-Knowledge and Attention-based Meta-Learning for Few-Shot Learning
- Meta-learning of Sequential Strategies
- Negative Margin Matters: Understanding Margin in Few-shot Classification
- TaskNorm: Rethinking Batch Normalization for Meta-Learning
- Incremental Few-Shot Learning with Attention Attractor Networks
- Edge-labeling Graph Neural Network for Few-shot Learning
- Meta-Learning for Low-Resource Neural Machine Translation
- Weight-Sharing Neural Architecture Search: A Battle to Shrink the Optimization Gap
- Meta-Baseline: Exploring Simple Meta-Learning for Few-Shot Learning
- Adversarial Meta-Learning
- Self-Supervised Learning For Few-Shot Image Classification
- Metalearned Neural Memory
- Deep Reinforcement Learning amidst Lifelong Non-Stationarity
- A Review and Comparison of AI Enhanced Side Channel Analysis
- Auto-Meta: Automated Gradient Based Meta Learner Search
- TGDM: Target Guided Dynamic Mixup for Cross-Domain Few-Shot Learning
- DPGN: Distribution Propagation Graph Network for Few-shot Learning
- Meta-DETR: Image-Level Few-Shot Object Detection with Inter-Class Correlation Exploitation
- Cross-Domain Few-Shot Learning by Representation Fusion
- Concept Learning through Deep Reinforcement Learning with Memory-Augmented Neural Networks
- Probabilistic Model-Agnostic Meta-Learning
- Weakly-supervised Compositional FeatureAggregation for Few-shot Recognition
- Image Deformation Meta-Networks for One-Shot Learning
- Large Margin Few-Shot Learning
- Adversarial Feature Hallucination Networks for Few-Shot Learning
- Metalearning with Hebbian Fast Weights
- Bongard-LOGO: A New Benchmark for Human-Level Concept Learning and Reasoning
- Task Augmentation by Rotating for Meta-Learning
- Penalty Method for Inversion-Free Deep Bilevel Optimization
- Sill-Net: Feature Augmentation with Separated Illumination Representation
- Infusing model predictive control into meta-reinforcement learning for mobile robots in dynamic environments
- Few-shot Learning with Meta Metric Learners
- Explanation-Guided Training for Cross-Domain Few-Shot Classification
- RelationNet2: Deep Comparison Columns for Few-Shot Learning
- Learning to Learn Variational Semantic Memory
- A Self Supervised StyleGAN for Image Annotation and Classification with Extremely Limited Labels
- BADGER: Learning to (Learn [Learning Algorithms] through Multi-Agent Communication)
- Behavior Self-Organization Supports Task Inference for Continual Robot Learning
- Differentiable plasticity: training plastic neural networks with backpropagation
- Revisiting Meta-Learning as Supervised Learning
- Induction Networks for Few-Shot Text Classification
- Sample-Efficient Reinforcement Learning via Counterfactual-Based Data Augmentation
- Towards Cross-Granularity Few-Shot Learning: Coarse-to-Fine Pseudo-Labeling with Visual-Semantic Meta-Embedding
- Prototype Rectification for Few-Shot Learning
- Watch, Try, Learn: Meta-Learning from Demonstrations and Reward
- Behavior Priors for Efficient Reinforcement Learning
- Few-shot Classification via Adaptive Attention
- On the Importance of Attention in Meta-Learning for Few-Shot Text Classification
- Toward Multimodal Model-Agnostic Meta-Learning
- Self-Attentional Credit Assignment for Transfer in Reinforcement Learning
- Reconciling meta-learning and continual learning with online mixtures of tasks
- Fixed-MAML for Few Shot Classification in Multilingual Speech Emotion Recognition
- Improving Generalization in Meta-learning via Task Augmentation
- A Policy Gradient Algorithm for Learning to Learn in Multiagent Reinforcement Learning
- Deep Metric Transfer for Label Propagation with Limited Annotated Data
- Meta-Learning across Meta-Tasks for Few-Shot Learning
- Meta-Learning with Fewer Tasks through Task Interpolation
- Learning and Planning with a Semantic Model
- Densely connected normalizing flows
- Scalable Bayesian Meta-Learning through Generalized Implicit Gradients
- Learning to learn via Self-Critique
- Offline Meta-Reinforcement Learning with Advantage Weighting
- Off-Dynamics Reinforcement Learning: Training for Transfer with Domain Classifiers
- Rapidly Adaptable Legged Robots via Evolutionary Meta-Learning
- Adaptive Posterior Learning: few-shot learning with a surprise-based memory module
- Improving Generalization via Scalable Neighborhood Component Analysis
- L2AE-D: Learning to Aggregate Embeddings for Few-shot Learning with Meta-level Dropout
- Context Meta-Reinforcement Learning via Neuromodulation
- Stochastic Prototype Embeddings
- Similarity R-C3D for Few-shot Temporal Activity Detection
- Uncertainty in Multitask Transfer Learning
- Weak-shot Fine-grained Classification via Similarity Transfer
- Variational Prototyping-Encoder: One-Shot Learning with Prototypical Images
- MetaCURE: Meta Reinforcement Learning with Empowerment-Driven Exploration
- MetaFun: Meta-Learning with Iterative Functional Updates
- An Investigation of Few-Shot Learning in Spoken Term Classification
- Revisiting Few-shot Activity Detection with Class Similarity Control
- Signal Transformer: Complex-valued Attention and Meta-Learning for Signal Recognition
- Learning from the Past: Continual Meta-Learning via Bayesian Graph Modeling
- Meta Learning Backpropagation And Improving It
- Online Structured Meta-learning
- Decoupling Exploration and Exploitation for Meta-Reinforcement Learning without Sacrifices
- Evolving Inborn Knowledge For Fast Adaptation in Dynamic POMDP Problems
- Curriculum in Gradient-Based Meta-Reinforcement Learning
- Knowledge as Priors: Cross-Modal Knowledge Generalization for Datasets without Superior Knowledge
- Bayesian Relational Memory for Semantic Visual Navigation
- TransMatch: A Transfer-Learning Scheme for Semi-Supervised Few-Shot Learning
- Meta-learning algorithms for Few-Shot Computer Vision
- Empirical Bayes Regret Minimization
- XtarNet: Learning to Extract Task-Adaptive Representation for Incremental Few-Shot Learning
- Task-similarity Aware Meta-learning through Nonparametric Kernel Regression
- Meta-Learning with Adaptive Hyperparameters
- AgileNet: Lightweight Dictionary-based Few-shot Learning
- MetaMixUp: Learning Adaptive Interpolation Policy of MixUp with Meta-Learning
- Few-Shot Knowledge Graph Completion
- Unsupervised Reinforcement Learning of Transferable Meta-Skills for Embodied Navigation
- Meta-Learning Initializations for Image Segmentation
- NDPNet: A novel non-linear data projection network for few-shot fine-grained image classification
- ECKPN: Explicit Class Knowledge Propagation Network for Transductive Few-shot Learning
- A Concise Review of Recent Few-shot Meta-learning Methods
- Adaptive Task Sampling for Meta-Learning
- Learning to Transfer: Unsupervised Meta Domain Translation
- Adaptive Cross-Modal Few-Shot Learning
- Differentiable Bandit Exploration
- Continual egocentric object recognition
- Modeling Taxi Drivers' Behaviour for the Next Destination Prediction
- Zero and Few Shot Learning with Semantic Feature Synthesis and Competitive Learning
- Embedding Adaptation is Still Needed for Few-Shot Learning
- Sequential Skip Prediction with Few-shot in Streamed Music Contents
- Self-Supervised Deep Visual Odometry with Online Adaptation
- Cross-Modulation Networks for Few-Shot Learning
- PAC-Bayes Bounds for Meta-learning with Data-Dependent Prior
- Meta-Learning Bandit Policies by Gradient Ascent
- Chameleon: Learning Model Initializations Across Tasks With Different Schemas
- Few-shot Sequence Learning with Transformers
- Meta-Model-Based Meta-Policy Optimization
- Meta Dropout: Learning to Perturb Features for Generalization
- Meta-Reinforcement Learning in Broad and Non-Parametric Environments
- Large-Scale Historical Watermark Recognition: dataset and a new consistency-based approach
- MetaFuse: A Pre-trained Fusion Model for Human Pose Estimation
- Human and Scene Motion Deblurring using Pseudo-blur Synthesizer
- AdarGCN: Adaptive Aggregation GCN for Few-Shot Learning
- Skill Transfer in Deep Reinforcement Learning under Morphological Heterogeneity
- On the Importance of Firth Bias Reduction in Few-Shot Classification
- Adversarial Attack across Datasets
- Discriminative Few-Shot Learning Based on Directional Statistics
- Few-Shot Learning with Intra-Class Knowledge Transfer
- Unsupervised Meta-Learning through Latent-Space Interpolation in Generative Models
- Attentive Feature Reuse for Multi Task Meta learning
- Adaptive Transformers in RL
- Are Fewer Labels Possible for Few-shot Learning?
- Bayes meets Bernstein at the Meta Level: an Analysis of Fast Rates in Meta-Learning with PAC-Bayes
- HetMAML: Task-Heterogeneous Model-Agnostic Meta-Learning for Few-Shot Learning Across Modalities
- Adaptive Adversarial Training for Meta Reinforcement Learning
- Bridging Text and Knowledge with Multi-Prototype Embedding for Few-Shot Relational Triple Extraction
- MELD: Meta-Reinforcement Learning from Images via Latent State Models
- Revisiting Mid-Level Patterns for Cross-Domain Few-Shot Recognition
- Physics-aware Spatiotemporal Modules with Auxiliary Tasks for Meta-Learning
- Zero-Shot Compositional Policy Learning via Language Grounding
- Variational Metric Scaling for Metric-Based Meta-Learning
- Decoder Choice Network for Meta-Learning
- Short-Term Stock Price-Trend Prediction Using Meta-Learning
- Few Shot Activity Recognition Using Variational Inference
- TOHAN: A One-step Approach towards Few-shot Hypothesis Adaptation
- Transductive Few-Shot Learning: Clustering is All You Need?
- VIABLE: Fast Adaptation via Backpropagating Learned Loss
- Hindsight Foresight Relabeling for Meta-Reinforcement Learning
- How to trust unlabeled data? Instance Credibility Inference for Few-Shot Learning
- Complementing Representation Deficiency in Few-shot Image Classification: A Meta-Learning Approach
- Reinforced Few-Shot Acquisition Function Learning for Bayesian Optimization
- Gaussian Process Meta Few-shot Classifier Learning via Linear Discriminant Laplace Approximation
- Multi-step Estimation for Gradient-based Meta-learning
- Pareto Self-Supervised Training for Few-Shot Learning
- Learning Associative Inference Using Fast Weight Memory
- One-shot Learning for Temporal Knowledge Graphs
- Learning to generate classifiers
- Testing the Genomic Bottleneck Hypothesis in Hebbian Meta-Learning
- Meta-Learning with Network Pruning
- An Optimization-Based Meta-Learning Model for MRI Reconstruction with Diverse Dataset
- On the Convergence Theory of Debiased Model-Agnostic Meta-Reinforcement Learning
- Provably Improved Context-Based Offline Meta-RL with Attention and Contrastive Learning
- Asymmetric Distribution Measure for Few-shot Learning
- Working Memory Graphs
- Task-Adaptive Clustering for Semi-Supervised Few-Shot Classification
- Progressive Cluster Purification for Transductive Few-shot Learning
- MLANE: Meta-Learning Based Adaptive Network Embedding
- Extended Few-Shot Learning: Exploiting Existing Resources for Novel Tasks
- Stabilizing Transformer-Based Action Sequence Generation For Q-Learning
- Large-Scale Meta-Learning with Continual Trajectory Shifting
- Few-shot Learning with Global Relatedness Decoupled-Distillation
- Task Affinity with Maximum Bipartite Matching in Few-Shot Learning
- Weak Novel Categories without Tears: A Survey on Weak-Shot Learning
- Tackling Early Sparse Gradients in Softmax Activation Using Leaky Squared Euclidean Distance
- Representation based and Attention augmented Meta learning
- Recomposing the Reinforcement Learning Building Blocks with Hypernetworks
- Meta-Reinforcement Learning for Heuristic Planning
- Enabling the Network to Surf the Internet
- Meta-Learner with Linear Nulling
- Memory-Augmented Relation Network for Few-Shot Learning
- MOTS: Multiple Object Tracking for General Categories Based On Few-Shot Method
- LFD-ProtoNet: Prototypical Network Based on Local Fisher Discriminant Analysis for Few-shot Learning
- Local Nonparametric Meta-Learning
- External-Memory Networks for Low-Shot Learning of Targets in Forward-Looking-Sonar Imagery
- Improve Unsupervised Pretraining for Few-label Transfer
- A Meta Reinforcement Learning-based Approach for Self-Adaptive System
- Adaptation-Agnostic Meta-Training
- Eden: A Unified Environment Framework for Booming Reinforcement Learning Algorithms
- Learning by Examples Based on Multi-level Optimization
- Detection of Adversarial Supports in Few-shot Classifiers Using Self-Similarity and Filtering
- Towards A Conceptually Simple Defensive Approach for Few-shot classifiers Against Adversarial Support Samples
- Meta-Learning Guarantees for Online Receding Horizon Learning Control
- Accelerating Distributed Online Meta-Learning via Multi-Agent Collaboration under Limited Communication
- HMRL: Hyper-Meta Learning for Sparse Reward Reinforcement Learning Problem
- Class Interference Regularization
- A Meta-Learning Control Algorithm with Provable Finite-Time Guarantees
- One-Shot Image Classification by Learning to Restore Prototypes
- MetalGAN: a Cluster-based Adaptive Training for Few-Shot Adversarial Colorization
- Shoestring: Graph-Based Semi-Supervised Learning with Severely Limited Labeled Data
- Unlocking the Full Potential of Small Data with Diverse Supervision
- Camera Distortion-aware 3D Human Pose Estimation in Video with Optimization-based Meta-Learning
- Transductive Learning for Textual Few-Shot Classification in API-based Embedding Models
- Augmented Bi-path Network for Few-shot Learning
- Top-Related Meta-Learning Method for Few-Shot Object Detection
- Meta-Learning with Adjoint Methods
- Meta R-CNN : Towards General Solver for Instance-level Few-shot Learning
- Task Attended Meta-Learning for Few-Shot Learning
- Mutual-Information Based Few-Shot Classification
- A Transductive Maximum Margin Classifier for Few-Shot Learning
- Cooperative Bi-path Metric for Few-shot Learning