Transient Non-Stationarity and Generalisation in Deep Reinforcement Learning
arXiv:2006.05826
Abstract
Non-stationarity can arise in Reinforcement Learning (RL) even in stationary environments. For example, most RL algorithms collect new data throughout training, using a non-stationary behaviour policy. Due to the transience of this non-stationarity, it is often not explicitly addressed in deep RL and a single neural network is continually updated. However, we find evidence that neural networks exhibit a memory effect where these transient non-stationarities can permanently impact the latent representation and adversely affect generalisation performance. Consequently, to improve generalisation of deep RL agents, we propose Iterated Relearning (ITER). ITER augments standard RL training by repeated knowledge transfer of the current policy into a freshly initialised network, which thereby experiences less non-stationarity during training. Experimentally, we show that ITER improves performance on the challenging generalisation benchmarks ProcGen and Multiroom.
References in corpus (9)
- Distilling the Knowledge in a Neural Network
- Measuring Catastrophic Forgetting in Neural Networks
- Leveraging Procedural Generation to Benchmark Reinforcement Learning
- Schema Networks: Zero-shot Transfer with a Generative Causal Model of Intuitive Physics
- Revisiting Fundamentals of Experience Replay
- Generalization in Reinforcement Learning with Selective Noise Injection and Information Bottleneck
- Distilling Policy Distillation
- Investigating Generalisation in Continuous Deep Reinforcement Learning
- Regularization Matters in Policy Optimization
Cited by in corpus (6)
- A Survey of Zero-shot Generalisation in Deep Reinforcement Learning
- Automated Reinforcement Learning (AutoRL): A Survey and Open Problems
- CORA: Benchmarks, Baselines, and Metrics as a Platform for Continual Reinforcement Learning Agents
- Augmented World Models Facilitate Zero-Shot Dynamics Generalization From a Single Offline Environment
- Reinforcement Learning by Guided Safe Exploration
- A study on the plasticity of neural networks