A Survey of Zero-shot Generalisation in Deep Reinforcement Learning
arXiv:2111.09794 · doi:10.1613/jair.1.14174
Abstract
The study of zero-shot generalisation (ZSG) in deep Reinforcement Learning (RL) aims to produce RL algorithms whose policies generalise well to novel unseen situations at deployment time, avoiding overfitting to their training environments. Tackling this is vital if we are to deploy reinforcement learning algorithms in real world scenarios, where the environment will be diverse, dynamic and unpredictable. This survey is an overview of this nascent field. We rely on a unifying formalism and terminology for discussing different ZSG problems, building upon previous works. We go on to categorise existing benchmarks for ZSG, as well as current methods for tackling these problems. Finally, we provide a critical discussion of the current state of the field, including recommendations for future work. Among other conclusions, we argue that taking a purely procedural content generation approach to benchmark design is not conducive to progress in ZSG, we suggest fast online adaptation and tackling RL-specific problems as some areas for future work on methods for ZSG, and we recommend building benchmarks in underexplored problem settings such as offline RL ZSG and reward-function variation.
JAIR version. Added formal definitions of ZSPT and related concepts, JAIR formatting, other small rewrites; https://www.jair.org/index.php/jair/article/view/14174
References in corpus (73)
- Improved Regularization of Convolutional Neural Networks with Cutout
- Domain Generalization: A Survey
- Solving Rubik's Cube with a Robot Hand
- Conservative Q-Learning for Offline Reinforcement Learning
- DeepMind Control Suite
- RL: Fast Reinforcement Learning via Slow Reinforcement Learning
- Robust Adversarial Reinforcement Learning
- robosuite: A Modular Simulation Framework and Benchmark for Robot Learning
- Transfer Learning in Deep Reinforcement Learning: A Survey
- Schema Networks: Zero-shot Transfer with a Generative Causal Model of Intuitive Physics
- Paired Open-Ended Trailblazer (POET): Endlessly Generating Increasingly Complex and Diverse Learning Environments and Their Solutions
- VariBAD: A Very Good Method for Bayes-Adaptive Deep RL via Meta-Learning
- Generalization in Reinforcement Learning with Selective Noise Injection and Information Bottleneck
- Open-Ended Learning Leads to Generally Capable Agents
- Environmental drivers of systematicity and generalization in a situated agent
- Can Autonomous Vehicles Identify, Recover From, and Adapt to Distribution Shifts?
- Phasic Policy Gradient
- Deep Online Learning via Meta-Learning: Continual Adaptation for Model-Based RL
- Network Randomization: A Simple Technique for Generalization in Deep Reinforcement Learning
- Improving Generalization in Reinforcement Learning with Mixture Regularization
- iGibson 1.0: a Simulation Environment for Interactive Tasks in Large Realistic Scenes
- Enhanced POET: Open-Ended Reinforcement Learning through Unbounded Invention of Learning Challenges and their Solutions
- Stabilizing Deep Q-Learning with ConvNets and Vision Transformers under Data Augmentation
- Investigating Generalisation in Continuous Deep Reinforcement Learning
- Is a Good Representation Sufficient for Sample Efficient Reinforcement Learning?
- Human Instruction-Following with Deep Reinforcement Learning via Transfer-Learning from Text
- CausalWorld: A Robotic Manipulation Benchmark for Causal Structure and Transfer Learning
- Towards Continual Reinforcement Learning: A Review and Perspectives
- Generalization of Reinforcement Learners with Working and Episodic Memory
- Contrastive Behavioral Similarity Embeddings for Generalization in Reinforcement Learning
- A Survey of Explainable Reinforcement Learning
- Online and Offline Reinforcement Learning by Planning with a Learned Model
- Emergent Complexity and Zero-shot Transfer via Unsupervised Environment Design
- Dynamics Generalization via Information Bottleneck in Deep Reinforcement Learning
- Observational Overfitting in Reinforcement Learning
- Generalization in Reinforcement Learning by Soft Data Augmentation
- SECANT: Self-Expert Cloning for Zero-Shot Generalization of Visual Policies
- The Distracting Control Suite -- A Challenging Benchmark for Reinforcement Learning from Pixels
- MiniHack the Planet: A Sandbox for Open-Ended Reinforcement Learning Research
- Procedural Content Generation: Better Benchmarks for Transfer Reinforcement Learning
- The Sensory Neuron as a Transformer: Permutation-Invariant Neural Networks for Reinforcement Learning
- Why Generalization in RL is Difficult: Epistemic POMDPs and Implicit Partial Observability
- Generalization to New Actions in Reinforcement Learning
- WordCraft: An Environment for Benchmarking Commonsense Agents
- CORA: Benchmarks, Baselines, and Metrics as a Platform for Continual Reinforcement Learning Agents
- Augmented World Models Facilitate Zero-Shot Dynamics Generalization From a Single Offline Environment
- Measuring Visual Generalization in Continuous Control from Pixels
- HALMA: Humanlike Abstraction Learning Meets Affordance in Rapid Problem Solving
- Cross-Trajectory Representation Learning for Zero-Shot Generalization in RL
- Trajectory-wise Multiple Choice Learning for Dynamics Generalization in Reinforcement Learning
- Domain Adversarial Reinforcement Learning
- Ask Your Humans: Using Human Instructions to Improve Generalization in Reinforcement Learning
- AdaRL: What, Where, and How to Adapt in Transfer Reinforcement Learning
- Toybox: A Suite of Environments for Experimental Evaluation of Deep Reinforcement Learning
- Generalization of Reinforcement Learning with Policy-Aware Adversarial Data Augmentation
- When Is Generalizable Reinforcement Learning Tractable?
- Reinforcement Learning with Videos: Combining Offline Observations with Interaction
- Robust Domain Randomised Reinforcement Learning through Peer-to-Peer Distillation
- A Survey of Exploration Methods in Reinforcement Learning
- Reasoning and Generalization in RL: A Tool Use Perspective
- Unsupervised Visual Attention and Invariance for Reinforcement Learning
- Neuro-algorithmic Policies enable Fast Combinatorial Generalization
- Procedural Generalization by Planning with Self-Supervised World Models
- Decoupling Value and Policy for Generalization in Reinforcement Learning
- Exploration in Approximate Hyper-State Space for Meta Reinforcement Learning
- Online Sparse Reinforcement Learning
- DRIBO: Robust Deep Reinforcement Learning via Multi-View Information Bottleneck
- CARL: A Benchmark for Contextual and Adaptive Reinforcement Learning
- Block Contextual MDPs for Continual Learning
- The Principle of Unchanged Optimality in Reinforcement Learning Generalization
- Sparse Attention Guided Dynamic Value Estimation for Single-Task Multi-Scene Reinforcement Learning
- Instance based Generalization in Reinforcement Learning
- Rogue-Gym: A New Challenge for Generalization in Reinforcement Learning
Cited by in corpus (11)
- Domain Generalization: A Survey
- State-of-the-art generalisation research in NLP: A taxonomy and review
- Generalization in Deep Reinforcement Learning for Robotic Navigation by Reward Shaping
- A Tutorial on Meta-Reinforcement Learning
- Multi-Agent Reinforcement Learning: Methods, Applications, Visionary Prospects, and Challenges
- Maximum diffusion reinforcement learning
- Structure in Deep Reinforcement Learning: A Survey and Open Problems
- GreenLight-Gym: Reinforcement learning benchmark environment for control of greenhouse production systems
- Sim-to-Real Transfer via a Style-Identified Cycle Consistent Generative Adversarial Network: Zero-Shot Deployment on Robotic Manipulators through Visual Domain Adaptation
- The RL Perceptron: Generalisation Dynamics of Policy Learning in High Dimensions
- Improving greenhouse fruit-production control by integrating reinforcement learning into short-horizon model predictive control