Improving Generalization in Meta-RL with Imaginary Tasks from Latent Dynamics Mixture
arXiv:2105.13524
Abstract
The generalization ability of most meta-reinforcement learning (meta-RL) methods is largely limited to test tasks that are sampled from the same distribution used to sample training tasks. To overcome the limitation, we propose Latent Dynamics Mixture (LDM) that trains a reinforcement learning agent with imaginary tasks generated from mixtures of learned latent dynamics. By training a policy on mixture tasks along with original training tasks, LDM allows the agent to prepare for unseen test tasks during training and prevents the agent from overfitting the training tasks. LDM significantly outperforms standard meta-RL methods in test returns on the gridworld navigation and MuJoCo tasks where we strictly separate the training task distribution and the test task distribution.
NeurIPS 2021
References in corpus (18)
- Efficient Off-Policy Meta-Reinforcement Learning via Probabilistic Context Variables
- Fast Context Adaptation via Meta-Learning
- Paired Open-Ended Trailblazer (POET): Endlessly Generating Increasingly Complex and Diverse Learning Environments and Their Solutions
- Self-Supervised Generalisation with Meta Auxiliary Learning
- VariBAD: A Very Good Method for Bayes-Adaptive Deep RL via Meta-Learning
- Improving Generalization in Meta Reinforcement Learning using Learned Objectives
- Meta reinforcement learning as task inference
- Generalization in Reinforcement Learning with Selective Noise Injection and Information Bottleneck
- Automatic Data Augmentation for Generalization in Deep Reinforcement Learning
- Enhanced POET: Open-Ended Reinforcement Learning through Unbounded Invention of Learning Challenges and their Solutions
- Meta-Q-Learning
- Virtual Class Enhanced Discriminative Embedding Learning
- Deep Reinforcement Learning amidst Lifelong Non-Stationarity
- Unsupervised Curricula for Visual Meta-Reinforcement Learning
- Meta-Reinforcement Learning Robust to Distributional Shift via Model Identification and Experience Relabeling
- Curriculum in Gradient-Based Meta-Reinforcement Learning
- Linear Representation Meta-Reinforcement Learning for Instant Adaptation
- Provably Efficient Model-based Policy Adaptation