Model-based Meta Reinforcement Learning using Graph Structured Surrogate Models
arXiv:2102.08291
Abstract
Reinforcement learning is a promising paradigm for solving sequential decision-making problems, but low data efficiency and weak generalization across tasks are bottlenecks in real-world applications. Model-based meta reinforcement learning addresses these issues by learning dynamics and leveraging knowledge from prior experience. In this paper, we take a closer look at this framework, and propose a new Thompson-sampling based approach that consists of a new model to identify task dynamics together with an amortized policy optimization step. We show that our model, called a graph structured surrogate model (GSSM), outperforms state-of-the-art methods in predicting environment dynamics. Additionally, our approach is able to obtain high returns, while allowing fast execution during deployment by avoiding test time policy gradient optimization.
References in corpus (7)
- Semi-Supervised Classification with Graph Convolutional Networks
- Mastering Chess and Shogi by Self-Play with a General Reinforcement Learning Algorithm
- Efficient Off-Policy Meta-Reinforcement Learning via Probabilistic Context Variables
- Exploring Model-based Planning with Policy Networks
- Attentive Neural Processes
- Context-aware Dynamics Model for Generalization in Model-Based Reinforcement Learning
- Meta-Learning surrogate models for sequential decision making