Data-Efficient Hierarchical Reinforcement Learning
arXiv:1805.08296
Abstract
Hierarchical reinforcement learning (HRL) is a promising approach to extend traditional reinforcement learning (RL) methods to solve more complex tasks. Yet, the majority of current HRL methods require careful task-specific design and on-policy training, making them difficult to apply in real-world scenarios. In this paper, we study how we can develop HRL algorithms that are general, in that they do not make onerous additional assumptions beyond standard RL algorithms, and efficient, in the sense that they can be used with modest numbers of interaction samples, making them suitable for real-world problems such as robotic control. For generality, we develop a scheme where lower-level controllers are supervised with goals that are learned and proposed automatically by the higher-level controllers. To address efficiency, we propose to use off-policy experience for both higher and lower-level training. This poses a considerable challenge, since changes to the lower-level behaviors change the action space for the higher-level policy, and we introduce an off-policy correction to remedy this challenge. This allows us to take advantage of recent advances in off-policy model-free RL to learn both higher- and lower-level policies using substantially fewer environment interactions than on-policy algorithms. We term the resulting HRL agent HIRO and find that it is generally applicable and highly sample-efficient. Our experiments show that HIRO can be used to learn highly complex behaviors for simulated robots, such as pushing objects and utilizing them to reach target locations, learning from only a few million samples, equivalent to a few days of real-time interaction. In comparisons with a number of prior HRL methods, we find that our approach substantially outperforms previous state-of-the-art techniques.
NIPS 2018
Cited by in corpus (109)
- Combinatorial Optimization by Graph Pointer Networks and Hierarchical Reinforcement Learning
- A Review of Robot Learning for Manipulation: Challenges, Representations, and Algorithms
- A survey on intrinsic motivation in reinforcement learning
- Dynamics-Aware Unsupervised Discovery of Skills
- RODE: Learning Roles to Decompose Multi-Agent Tasks
- Accelerating Reinforcement Learning for Reaching using Continuous Curriculum Learning
- An information-theoretic perspective on intrinsic motivation in reinforcement learning: a survey
- Why Does Hierarchy (Sometimes) Work So Well in Reinforcement Learning?
- Deployment-Efficient Reinforcement Learning via Model-Based Offline Optimization
- CURIOUS: Intrinsically Motivated Modular Multi-Goal Reinforcement Learning
- Search on the Replay Buffer: Bridging Planning and Reinforcement Learning
- Hierarchical Deep Multiagent Reinforcement Learning with Temporal Abstraction
- From Machine Learning to Robotics: Challenges and Opportunities for Embodied Intelligence
- Learning a Contact-Adaptive Controller for Robust, Efficient Legged Locomotion
- Mapping State Space using Landmarks for Universal Goal Reaching
- Generating Adjacency-Constrained Subgoals in Hierarchical Reinforcement Learning
- Learning to Set Waypoints for Audio-Visual Navigation
- Parrot: Data-Driven Behavioral Priors for Reinforcement Learning
- Hierarchical Reinforcement Learning By Discovering Intrinsic Options
- Learning Functionally Decomposed Hierarchies for Continuous Control Tasks with Path Planning
- Asymmetric self-play for automatic goal discovery in robotic manipulation
- AutoQ: Automated Kernel-Wise Neural Network Quantization
- Learning Compositional Neural Programs with Recursive Tree Search and Planning
- Broadly-Exploring, Local-Policy Trees for Long-Horizon Task Planning
- Landmark-Guided Subgoal Generation in Hierarchical Reinforcement Learning
- From Pixels to Legs: Hierarchical Learning of Quadruped Locomotion
- State Alignment-based Imitation Learning
- Structure in Deep Reinforcement Learning: A Survey and Open Problems
- Reinforcement Learning for Multi-Product Multi-Node Inventory Management in Supply Chains
- Explore, Discover and Learn: Unsupervised Discovery of State-Covering Skills
- OPAL: Offline Primitive Discovery for Accelerating Offline Reinforcement Learning
- Contextual Imagined Goals for Self-Supervised Robotic Learning
- Recent Advances in Leveraging Human Guidance for Sequential Decision-Making Tasks
- Behavior Priors for Efficient Reinforcement Learning
- D2RL: Deep Dense Architectures in Reinforcement Learning
- Sub-policy Adaptation for Hierarchical Reinforcement Learning
- Relay Policy Learning: Solving Long-Horizon Tasks via Imitation and Reinforcement Learning
- HRL4IN: Hierarchical Reinforcement Learning for Interactive Navigation with Mobile Manipulators
- Reset-Free Lifelong Learning with Skill-Space Planning
- Hierarchical Reinforcement Learning for Quadruped Locomotion
- Hierarchical Policy Learning is Sensitive to Goal Space Design
- Composing Task-Agnostic Policies with Deep Reinforcement Learning
- Adversarial Option-Aware Hierarchical Imitation Learning
- Directed Exploration for Reinforcement Learning
- Disentangled Skill Embeddings for Reinforcement Learning
- Solving Compositional Reinforcement Learning Problems via Task Reduction
- Learning Hierarchical Teaching Policies for Cooperative Agents
- Example-Driven Model-Based Reinforcement Learning for Solving Long-Horizon Visuomotor Tasks
- DREAM Architecture: a Developmental Approach to Open-Ended Learning in Robotics
- TAAC: Temporally Abstract Actor-Critic for Continuous Control
- TRAIL: Near-Optimal Imitation Learning with Suboptimal Data
- Accelerating Robotic Reinforcement Learning via Parameterized Action Primitives
- Resolving Spurious Correlations in Causal Models of Environments via Interventions
- Hierarchical Reinforcement Learning for Relay Selection and Power Optimization in Two-Hop Cooperative Relay Network
- Bi-level Off-policy Reinforcement Learning for Volt/VAR Control Involving Continuous and Discrete Devices
- Following Instructions by Imagining and Reaching Visual Goals
- Learning and Exploiting Multiple Subgoals for Fast Exploration in Hierarchical Reinforcement Learning
- Generalized Decision Transformer for Offline Hindsight Information Matching
- Predictive Coding for Boosting Deep Reinforcement Learning with Sparse Rewards
- Provable Hierarchical Imitation Learning via EM
- Skill Transfer in Deep Reinforcement Learning under Morphological Heterogeneity
- Towards Autonomous Pipeline Inspection with Hierarchical Reinforcement Learning
- Outcome-Driven Reinforcement Learning via Variational Inference
- Learning Representations in Reinforcement Learning:An Information Bottleneck Approach
- Learning Task Decomposition with Ordered Memory Policy Network
- Hierarchies of Planning and Reinforcement Learning for Robot Navigation
- Hierarchically Integrated Models: Learning to Navigate from Heterogeneous Robots
- Online Baum-Welch algorithm for Hierarchical Imitation Learning
- Correcting Experience Replay for Multi-Agent Communication
- Planning in Learned Latent Action Spaces for Generalizable Legged Locomotion
- Inter-Level Cooperation in Hierarchical Reinforcement Learning
- Learning to Solve a Rubik's Cube with a Dexterous Hand
- Control What You Can: Intrinsically Motivated Task-Planning Agent
- Exploration via Hindsight Goal Generation
- Avoiding Tampering Incentives in Deep RL via Decoupled Approval
- Learning Compositional Neural Programs for Continuous Control
- Density-based Curriculum for Multi-goal Reinforcement Learning with Sparse Rewards
- A Deep Reinforcement Learning Architecture for Multi-stage Optimal Control
- Goal-conditioned Batch Reinforcement Learning for Rotation Invariant Locomotion
- Abstract Value Iteration for Hierarchical Reinforcement Learning
- Beyond Tabula-Rasa: a Modular Reinforcement Learning Approach for Physically Embedded 3D Sokoban
- From semantics to execution: Integrating action planning with reinforcement learning for robotic causal problem-solving
- OCEAN: Online Task Inference for Compositional Tasks with Context Adaptation
- HAC Explore: Accelerating Exploration with Hierarchical Reinforcement Learning
- Learning the Solution Manifold in Optimization and Its Application in Motion Planning
- TTR-Based Reward for Reinforcement Learning with Implicit Model Priors
- Hierarchical Skills for Efficient Exploration
- Self-supervised Reinforcement Learning with Independently Controllable Subgoals
- Efficient Robotic Object Search via HIEM: Hierarchical Policy Learning with Intrinsic-Extrinsic Modeling
- A survey of benchmarking frameworks for reinforcement learning
- C-Learning: Horizon-Aware Cumulative Accessibility Estimation
- Spatially and Seamlessly Hierarchical Reinforcement Learning for State Space and Policy space in Autonomous Driving
- Training Transition Policies via Distribution Matching for Complex Tasks
- Interviewer-Candidate Role Play: Towards Developing Real-World NLP Systems
- Hierarchical Neural Dynamic Policies
- Generalization in Text-based Games via Hierarchical Reinforcement Learning
- Feudal Reinforcement Learning by Reading Manuals
- Provable Hierarchy-Based Meta-Reinforcement Learning
- Distilling a Hierarchical Policy for Planning and Control via Representation and Reinforcement Learning
- Active Hierarchical Imitation and Reinforcement Learning
- Playing Atari Ball Games with Hierarchical Reinforcement Learning
- Computational principles of intelligence: learning and reasoning with neural networks
- Hierarchical and Partially Observable Goal-driven Policy Learning with Goals Relational Graph
- A New Framework for Machine Intelligence: Concepts and Prototype
- Developing cooperative policies for multi-stage tasks
- Scalable, Decentralized Multi-Agent Reinforcement Learning Methods Inspired by Stigmergy and Ant Colonies
- Hierarchical Policies for Cluttered-Scene Grasping with Latent Plans
- Hierarchical Reinforcement Learning with Optimal Level Synchronization Based on Flow-Based Deep Generative Model
- HILONet: Hierarchical Imitation Learning from Non-Aligned Observations