Latent Space Policies for Hierarchical Reinforcement Learning
arXiv:1804.02808
Abstract
We address the problem of learning hierarchical deep neural network policies for reinforcement learning. In contrast to methods that explicitly restrict or cripple lower layers of a hierarchy to force them to use higher-level modulating signals, each layer in our framework is trained to directly solve the task, but acquires a range of diverse strategies via a maximum entropy reinforcement learning objective. Each layer is also augmented with latent random variables, which are sampled from a prior distribution during the training of that layer. The maximum entropy objective causes these latent variables to be incorporated into the layer's policy, and the higher level layer can directly control the behavior of the lower layer through this latent space. Furthermore, by constraining the mapping from latent variables to actions to be invertible, higher layers retain full expressivity: neither the higher layers nor the lower layers are constrained in their behavior. Our experimental evaluation demonstrates that we can improve on the performance of single-layer policies on standard benchmark tasks simply by adding additional layers, and that our method can solve more complex sparse-reward tasks by learning higher-level policies on top of high-entropy skills optimized for simple low-level objectives.
ICML 2018; Videos: https://sites.google.com/view/latent-space-deep-rl Code: https://github.com/haarnoja/sac
References in corpus (2)
Cited by in corpus (31)
- Discrete and Continuous Action Representation for Practical RL in Video Games
- Parrot: Data-Driven Behavioral Priors for Reinforcement Learning
- Reinforcement Learning with Competitive Ensembles of Information-Constrained Primitives
- Boosting Trust Region Policy Optimization by Normalizing Flows Policy
- Structure in Deep Reinforcement Learning: A Survey and Open Problems
- Dynamics-aware Embeddings
- Discretizing Continuous Action Space for On-Policy Optimization
- Learning Goal Embeddings via Self-Play for Hierarchical Reinforcement Learning
- Leveraging exploration in off-policy algorithms via normalizing flows
- Hierarchical Reinforcement Learning for Quadruped Locomotion
- SUMO: Unbiased Estimation of Log Marginal Probability for Latent Variable Models
- Adapting Behaviour for Learning Progress
- Reinforced Wasserstein Training for Severity-Aware Semantic Segmentation in Autonomous Driving
- Greedy Hierarchical Variational Autoencoders for Large-Scale Video Prediction
- VFunc: a Deep Generative Model for Functions
- Implicit Policy for Reinforcement Learning
- Hindsight Expectation Maximization for Goal-conditioned Reinforcement Learning
- Catalyst.RL: A Distributed Framework for Reproducible RL Research
- Quinoa: a Q-function You Infer Normalized Over Actions
- Distributed Soft Actor-Critic with Multivariate Reward Representation and Knowledge Distillation
- Skill Discovery of Coordination in Multi-agent Reinforcement Learning
- Diversity-Driven Extensible Hierarchical Reinforcement Learning
- Energy-based Surprise Minimization for Multi-Agent Value Factorization
- Sample Efficient Ensemble Learning with Catalyst.RL
- Feudal Steering: Hierarchical Learning for Steering Angle Prediction
- Generative Actor-Critic: An Off-policy Algorithm Using the Push-forward Model
- Neural Embedding for Physical Manipulations
- From proprioception to long-horizon planning in novel environments: A hierarchical RL model
- Developing cooperative policies for multi-stage tasks
- Transfer Learning by Modeling a Distribution over Policies
- A Bayesian Approach to Identifying Representational Errors