Multi-Level Discovery of Deep Options
arXiv:1703.08294
Abstract
Augmenting an agent's control with useful higher-level behaviors called options can greatly reduce the sample complexity of reinforcement learning, but manually designing options is infeasible in high-dimensional and abstract state spaces. While recent work has proposed several techniques for automated option discovery, they do not scale to multi-level hierarchies and to expressive representations such as deep networks. We present Discovery of Deep Options (DDO), a policy-gradient algorithm that discovers parametrized options from a set of demonstration trajectories, and can be used recursively to discover additional levels of the hierarchy. The scalability of our approach to multi-level hierarchies stems from the decoupling of low-level option discovery from high-level meta-control policy learning, facilitated by under-parametrization of the high level. We demonstrate that using the discovered options to augment the action space of Deep Q-Network agents can accelerate learning by guiding exploration in tasks where random actions are unlikely to reach valuable states. We show that DDO is effective in adding options that accelerate learning in 4 out of 5 Atari RAM environments chosen in our experiments. We also show that DDO can discover structure in robot-assisted surgical videos and kinematics that match expert annotation with 72% accuracy.
References in corpus (4)
Cited by in corpus (36)
- Gesture Recognition in Robotic Surgery: a Review
- Variational End-to-End Navigation and Localization
- Variational Option Discovery Algorithms
- Deep Hierarchical Reinforcement Learning Algorithm in Partially Observable Markov Decision Processes
- Learning Abstract Options
- Search on the Replay Buffer: Bridging Planning and Reinforcement Learning
- TACO: Learning Task Decomposition via Temporal Alignment for Control
- Multi-Modal Imitation Learning from Unstructured Demonstrations using Generative Adversarial Nets
- Parrot: Data-Driven Behavioral Priors for Reinforcement Learning
- Broadly-Exploring, Local-Policy Trees for Long-Horizon Task Planning
- Exploiting Hierarchy for Learning and Transfer in KL-regularized RL
- From Pixels to Legs: Hierarchical Learning of Quadruped Locomotion
- Learning Robust Bed Making using Deep Imitation Learning with DART
- Explore, Discover and Learn: Unsupervised Discovery of State-Covering Skills
- Behavior Priors for Efficient Reinforcement Learning
- Modeling Long-horizon Tasks as Sequential Interaction Landscapes
- Relay Policy Learning: Solving Long-Horizon Tasks via Imitation and Reinforcement Learning
- CompILE: Compositional Imitation Learning and Execution
- DDCO: Discovery of Deep Continuous Options for Robot Learning from Demonstrations
- Data-efficient Hindsight Off-policy Option Learning
- SOAC: The Soft Option Actor-Critic Architecture
- Idiosyncrasies and challenges of data driven learning in electronic trading
- DREAM Architecture: a Developmental Approach to Open-Ended Learning in Robotics
- TRAIL: Near-Optimal Imitation Learning with Suboptimal Data
- Self-organization of action hierarchy and compositionality by reinforcement learning with recurrent neural networks
- Provable Hierarchical Imitation Learning via EM
- Value Function Spaces: Skill-Centric State Abstractions for Long-Horizon Reasoning
- Diversity-Driven Extensible Hierarchical Reinforcement Learning
- Learning Task Decomposition with Ordered Memory Policy Network
- Goal Kernel Planning: Linearly-Solvable Non-Markovian Policies for Logical Tasks with Goal-Conditioned Options
- Online Baum-Welch algorithm for Hierarchical Imitation Learning
- Hierarchical Representation Learning for Markov Decision Processes
- PLOTS: Procedure Learning from Observations using Subtask Structure
- Bottom-Up Skill Discovery from Unsegmented Demonstrations for Long-Horizon Robot Manipulation
- Reinforcement Learning via Reasoning from Demonstration
- Discovering hierarchies using Imitation Learning from hierarchy aware policies