Meta Learning Shared Hierarchies
arXiv:1710.09767
Abstract
We develop a metalearning approach for learning hierarchically structured policies, improving sample efficiency on unseen tasks through the use of shared primitives---policies that are executed for large numbers of timesteps. Specifically, a set of primitives are shared within a distribution of tasks, and are switched between by task-specific policies. We provide a concrete metric for measuring the strength of such hierarchies, leading to an optimization problem for quickly reaching high reward on unseen tasks. We then present an algorithm to solve this problem end-to-end through the use of any off-the-shelf reinforcement learning method, by repeatedly sampling new tasks and resetting task-specific policies. We successfully discover meaningful motor primitives for the directional movement of four-legged robots, solely by interacting with distributions of mazes. We also demonstrate the transferability of primitives to solve long-timescale sparse-reward obstacle courses, and we enable 3D humanoid robots to robustly walk and crawl with the same policy.
References in corpus (1)
Cited by in corpus (18)
- Search on the Replay Buffer: Bridging Planning and Reinforcement Learning
- Transfer in Deep Reinforcement Learning Using Successor Features and Generalised Policy Improvement
- Reinforcement Learning with Competitive Ensembles of Information-Constrained Primitives
- Learning to Compose Skills
- Generalized Hindsight for Reinforcement Learning
- Hierarchical Reinforcement Learning for Multi-agent MOBA Game
- Hierarchical Policy Learning is Sensitive to Goal Space Design
- Hierarchical Reinforcement Learning for Quadruped Locomotion
- Learning Generalizable Locomotion Skills with Hierarchical Reinforcement Learning
- Contextualizing Enhances Gradient Based Meta Learning
- Optimal Options for Multi-Task Reinforcement Learning Under Time Constraints
- Continual and Multi-task Reinforcement Learning With Shared Episodic Memory
- Learning to Cope with Adversarial Attacks
- A Deep Reinforcement Learning Architecture for Multi-stage Optimal Control
- Reinforcement Learning Experience Reuse with Policy Residual Representation
- Temporal-adaptive Hierarchical Reinforcement Learning
- Developing cooperative policies for multi-stage tasks
- On mechanisms for transfer using landmark value functions in multi-task lifelong reinforcement learning