Skill Transfer in Deep Reinforcement Learning under Morphological Heterogeneity
arXiv:1908.05265
Abstract
Transfer learning methods for reinforcement learning (RL) domains facilitate the acquisition of new skills using previously acquired knowledge. The vast majority of existing approaches assume that the agents have the same design, e.g. same shape and action spaces. In this paper we address the problem of transferring previously acquired skills amongst morphologically different agents (MDAs). For instance, assuming that a bipedal agent has been trained to move forward, could this skill be transferred on to a one-leg hopper so as to make its training process for the same task more sample efficient? We frame this problem as one of subspace learning whereby we aim to infer latent factors representing the control mechanism that is common between MDAs. We propose a novel paired variational encoder-decoder model, PVED, that disentangles the control of MDAs into shared and agent-specific factors. The shared factors are then leveraged for skill transfer using RL. Theoretically, we derive a theorem indicating how the performance of PVED depends on the shared factors and agent morphologies. Experimentally, PVED has been extensively validated on four MuJoCo environments. We demonstrate its performance compared to a state-of-the-art approach and several ablation cases, visualize and interpret the hidden factors, and identify avenues for future improvements.
References in corpus (16)
- Recurrent World Models Facilitate Policy Evolution
- Variational Autoencoder for Deep Learning of Images, Labels and Captions
- Visual Reinforcement Learning with Imagined Goals
- FeUdal Networks for Hierarchical Reinforcement Learning
- Disentangling factors of variation in deep representations using adversarial training
- Data-Efficient Hierarchical Reinforcement Learning
- Zero-Shot Task Generalization with Multi-Task Deep Reinforcement Learning
- Learning Invariant Feature Spaces to Transfer Skills with Reinforcement Learning
- Policy Transfer with Strategy Optimization
- Efficient Model-Based Deep Reinforcement Learning with Variational State Tabulation
- Hardware Conditioned Policies for Multi-Robot Transfer Learning
- Mix&Match - Agent Curricula for Reinforcement Learning
- Learning Discrete and Continuous Factors of Data via Alternating Disentanglement
- Exploiting Hierarchy for Learning and Transfer in KL-regularized RL
- Synthesized Policies for Transfer and Adaptation across Tasks and Environments
- Transfer Learning From Synthetic To Real Images Using Variational Autoencoders For Precise Position Detection