Attentive Neural Processes
arXiv:1901.05761
Abstract
Neural Processes (NPs) (Garnelo et al 2018a;b) approach regression by learning to map a context set of observed input-output pairs to a distribution over regression functions. Each function models the distribution of the output given an input, conditioned on the context. NPs have the benefit of fitting observed data efficiently with linear complexity in the number of context input-output pairs, and can learn a wide family of conditional distributions; they learn predictive distributions conditioned on context sets of arbitrary size. Nonetheless, we show that NPs suffer a fundamental drawback of underfitting, giving inaccurate predictions at the inputs of the observed data they condition on. We address this issue by incorporating attention into NPs, allowing each input location to attend to the relevant context points for the prediction. We show that this greatly improves the accuracy of predictions, results in noticeably faster training, and expands the range of functions that can be modelled.
References in corpus (9)
- Auto-Encoding Variational Bayes
- Prototypical Networks for Few-shot Learning
- One-shot Learning with Memory-Augmented Neural Networks
- Few-shot Autoregressive Density Estimation: Towards Learning to Learn Distributions
- The Variational Homoencoder: Learning to learn high capacity generative models from few examples
- Learning models for visual 3D localization with implicit mapping
- Consistent Generative Query Networks
- Variational Implicit Processes
- VFunc: a Deep Generative Model for Functions
Cited by in corpus (20)
- Graph Element Networks: adaptive, structured computation and memory
- Adaptive Deep Kernel Learning
- Probabilistic Trajectory Prediction for Autonomous Vehicles with Attentive Recurrent Neural Process
- Meta-Amortized Variational Inference and Learning
- Noise Contrastive Meta-Learning for Conditional Density Estimation using Kernel Mean Embeddings
- Context-Aware Safe Reinforcement Learning for Non-Stationary Environments
- Neural Process for Black-Box Model Optimization Under Bayesian Framework
- Contrastive Neural Processes for Self-Supervised Learning
- Function Contrastive Learning of Transferable Meta-Representations
- Energy-Based Processes for Exchangeable Data
- Model-based Meta Reinforcement Learning using Graph Structured Surrogate Models
- Reinforcement Learning for Robotics and Control with Active Uncertainty Reduction
- Mini-Batch Consistent Slot Set Encoder for Scalable Set Encoding
- Meta Learning as Bayes Risk Minimization
- Robustifying Sequential Neural Processes
- Inferring Black Hole Properties from Astronomical Multivariate Time Series with Bayesian Attentive Neural Processes
- Meta-Learning for Koopman Spectral Analysis with Short Time-series
- OR-Net: Pointwise Relational Inference for Data Completion under Partial Observation
- MFPC-Net: Multi-fidelity Physics-Constrained Neural Process
- CAZSL: Zero-Shot Regression for Pushing Models by Generalizing Through Context