Generalization in Transfer Learning
arXiv:1909.01331 · doi:10.1017/S0263574722000625
Abstract
Agents trained with deep reinforcement learning algorithms are capable of performing highly complex tasks including locomotion in continuous environments. We investigate transferring the learning acquired in one task to a set of previously unseen tasks. Generalization and overfitting in deep reinforcement learning are not commonly addressed in current transfer learning research. Conducting a comparative analysis without an intermediate regularization step results in underperforming benchmarks and inaccurate algorithm comparisons due to rudimentary assessments. In this study, we propose regularization techniques in deep reinforcement learning for continuous control through the application of sample elimination, early stopping and maximum entropy regularized adversarial learning. First, the importance of the inclusion of training iteration number to the hyperparameters in deep transfer reinforcement learning will be discussed. Because source task performance is not indicative of the generalization capacity of the algorithm, we start by acknowledging the training iteration number as a hyperparameter. In line with this, we introduce an additional step of resorting to earlier snapshots of policy parameters to prevent overfitting to the source task. Then, to generate robust policies, we discard the samples that lead to overfitting via a method we call strict clipping. Furthermore, we increase the generalization capacity in widely used transfer learning benchmarks by using maximum entropy regularization, different critic methods, and curriculum learning in an adversarial setup. Subsequently, we propose maximum entropy adversarial reinforcement learning to increase the domain randomization. Finally, we evaluate the robustness of these methods on simulated robots in target environments where the morphology of the robot, gravity, and tangential friction coefficient of the environment are altered.
23 pages, 36 figures
References in corpus (14)
- Model-Agnostic Meta-Learning for Fast Adaptation of Deep Networks
- Soft Actor-Critic: Off-Policy Maximum Entropy Deep Reinforcement Learning with a Stochastic Actor
- Learning agile and dynamic motor skills for legged robots
- Robust Adversarial Reinforcement Learning
- A Study on Overfitting in Deep Reinforcement Learning
- Quantifying Generalization in Reinforcement Learning
- Continuous Adaptation via Meta-Learning in Nonstationary and Competitive Environments
- EPOpt: Learning Robust Neural Network Policies Using Model Ensembles
- Emergent Complexity via Multi-Agent Competition
- Learning to Walk in the Real World with Minimal Human Effort
- Iterative Reinforcement Learning Based Design of Dynamic Locomotion Skills for Cassie
- Investigating Generalisation in Continuous Deep Reinforcement Learning
- Clipped Action Policy Gradient
- Benchmark Environments for Multitask Learning in Continuous Domains
Cited by in corpus (3)
- Diffusion Policies for Out-of-Distribution Generalization in Offline Reinforcement Learning
- Bidirectional Progressive Neural Networks with Episodic Return Progress for Emergent Task Sequencing and Robotic Skill Transfer
- Unsupervised Meta-Testing with Conditional Neural Processes for Hybrid Meta-Reinforcement Learning