Hindsight Experience Replay
arXiv:1707.01495
Abstract
Dealing with sparse rewards is one of the biggest challenges in Reinforcement Learning (RL). We present a novel technique called Hindsight Experience Replay which allows sample-efficient learning from rewards which are sparse and binary and therefore avoid the need for complicated reward engineering. It can be combined with an arbitrary off-policy RL algorithm and may be seen as a form of implicit curriculum. We demonstrate our approach on the task of manipulating objects with a robotic arm. In particular, we run experiments on three different tasks: pushing, sliding, and pick-and-place, in each case using only binary rewards indicating whether or not the task is completed. Our ablation studies show that Hindsight Experience Replay is a crucial ingredient which makes training possible in these challenging environments. We show that our policies trained on a physics simulation can be deployed on a physical robot and successfully complete the task.
References in corpus (6)
- Domain Randomization for Transferring Deep Neural Networks from Simulation to the Real World
- Automatic Goal Generation for Reinforcement Learning Agents
- Data-efficient Deep Reinforcement Learning for Dexterous Manipulation
- Discrete Sequential Prediction of Continuous Actions for Deep RL
- Automated Curriculum Learning for Neural Networks
- First Experiments with PowerPlay
Cited by in corpus (196)
- A Survey and Critique of Multiagent Deep Reinforcement Learning
- How to Train Your Robot with Deep Reinforcement Learning; Lessons We've Learned
- Data-Efficient Hierarchical Reinforcement Learning
- Reinforcement Learning with Augmented Data
- Maximum a Posteriori Policy Optimisation
- Reward Machines: Exploiting Reward Function Structure in Reinforcement Learning
- A Deeper Look at Experience Replay
- Ablation Studies in Artificial Neural Networks
- Temporal Difference Models: Model-Free Deep RL for Model-Based Control
- Emergent Complexity via Multi-Agent Competition
- Automatic Goal Generation for Reinforcement Learning Agents
- A Theoretical Analysis of Deep Q-Learning
- Illuminating Generalization in Deep Reinforcement Learning through Procedural Level Generation
- Grasp2Vec: Learning Object Representations from Self-Supervised Grasping
- A Review of Robot Learning for Manipulation: Challenges, Representations, and Algorithms
- Universal Planning Networks
- GEP-PG: Decoupling Exploration and Exploitation in Deep Reinforcement Learning Algorithms
- Dynamics-Aware Unsupervised Discovery of Skills
- t-Soft Update of Target Network for Deep Reinforcement Learning
- Language as a Cognitive Tool to Imagine Goals in Curiosity-Driven Exploration
- Unsupervised Meta-Learning for Reinforcement Learning
- Accelerating Reinforcement Learning for Reaching using Continuous Curriculum Learning
- Dealing with Sparse Rewards in Reinforcement Learning
- Causal Induction from Visual Observations for Goal Directed Tasks
- Reinforcement Learning with Prototypical Representations
- CURIOUS: Intrinsically Motivated Modular Multi-Goal Reinforcement Learning
- Hierarchical Reinforcement Learning with Hindsight
- Learning offline: memory replay in biological and artificial reinforcement learning
- Composable Planning with Attributes
- Hindsight policy gradients
- Unicorn: Continual Learning with a Universal, Off-policy Agent
- DeepRacer: Educational Autonomous Racing Platform for Experimentation with Sim2Real Reinforcement Learning
- Computational Theories of Curiosity-Driven Learning
- Placeto: Learning Generalizable Device Placement Algorithms for Distributed Machine Learning
- Rewriting History with Inverse RL: Hindsight Inference for Policy Improvement
- ManiSkill: Generalizable Manipulation Skill Benchmark with Large-Scale Demonstrations
- A Geometric Perspective on Optimal Representations for Reinforcement Learning
- Deep Reinforcement Learning Methods for Structure-Guided Processing Path Optimization
- Energy-Based Hindsight Experience Prioritization
- VAT-Mart: Learning Visual Action Trajectory Proposals for Manipulating 3D ARTiculated Objects
- Learning Functionally Decomposed Hierarchies for Continuous Control Tasks with Path Planning
- Universal Successor Features Approximators
- Learning 6-DoF Grasping and Pick-Place Using Attention Focus
- Learning Latent Plans from Play
- Evolving Rewards to Automate Reinforcement Learning
- World Model as a Graph: Learning Latent Landmarks for Planning
- Replay in Deep Learning: Current Approaches and Missing Biological Elements
- Variational Dynamic for Self-Supervised Exploration in Deep Reinforcement Learning
- Improving Target-driven Visual Navigation with Attention on 3D Spatial Relationships
- ScreenerNet: Learning Self-Paced Curriculum for Deep Neural Networks
- ACTRCE: Augmenting Experience via Teacher's Advice For Multi-Goal Reinforcement Learning
- ARCHER: Aggressive Rewards to Counter bias in Hindsight Experience Replay
- Reinforcement Learning of Active Vision for Manipulating Objects under Occlusions
- Optimizing Throughput Performance in Distributed MIMO Wi-Fi Networks using Deep Reinforcement Learning
- PlanGAN: Model-based Planning With Sparse Rewards and Multiple Goals
- An Open-Source Multi-Goal Reinforcement Learning Environment for Robotic Manipulation with Pybullet
- Training Agents using Upside-Down Reinforcement Learning
- Robust Policies via Mid-Level Visual Representations: An Experimental Study in Manipulation and Navigation
- Improving Robot Dual-System Motor Learning with Intrinsically Motivated Meta-Control and Latent-Space Experience Imagination
- Experience Replay Optimization
- Actionable Models: Unsupervised Offline Reinforcement Learning of Robotic Skills
- Reinforcement Learning on Web Interfaces Using Workflow-Guided Exploration
- Hierarchical Foresight: Self-Supervised Learning of Long-Horizon Tasks via Visual Subgoal Generation
- Explore, Discover and Learn: Unsupervised Discovery of State-Covering Skills
- Goal-Auxiliary Actor-Critic for 6D Robotic Grasping with Point Clouds
- What Did You Think Would Happen? Explaining Agent Behaviour Through Intended Outcomes
- Striving for Simplicity and Performance in Off-Policy DRL: Output Normalization and Non-Uniform Sampling
- Time Reversal as Self-Supervision
- Reinforcement Learning Upside Down: Don't Predict Rewards -- Just Map Them to Actions
- Learning Language-Conditioned Robot Behavior from Offline Data and Crowd-Sourced Annotation
- A Survey of Reinforcement Learning Techniques: Strategies, Recent Development, and Future Directions
- Learning Obstacle Representations for Neural Motion Planning
- A Boolean Task Algebra for Reinforcement Learning
- Behavior From the Void: Unsupervised Active Pre-Training
- Options as responses: Grounding behavioural hierarchies in multi-agent RL
- Divide-and-Conquer Reinforcement Learning
- Motion Perception in Reinforcement Learning with Dynamic Objects
- On the Weaknesses of Reinforcement Learning for Neural Machine Translation
- Dynamic Experience Replay
- Multi-Task Domain Adaptation for Deep Learning of Instance Grasping from Simulation
- Self-Paced Contextual Reinforcement Learning
- Adapting Behaviour via Intrinsic Reward: A Survey and Empirical Study
- Differentiable Logic Machines
- Learning Manipulation States and Actions for Efficient Non-prehensile Rearrangement Planning
- Data-efficient Hindsight Off-policy Option Learning
- ROLL: Visual Self-Supervised Reinforcement Learning with Object Reasoning
- Hindsight Trust Region Policy Optimization
- Sub-Goal Trees -- a Framework for Goal-Based Reinforcement Learning
- Offline Meta-Reinforcement Learning with Online Self-Supervision
- Deep Reinforcement Learning to Acquire Navigation Skills for Wheel-Legged Robots in Complex Environments
- The Effect of Multi-step Methods on Overestimation in Deep Reinforcement Learning
- ELLA: Exploration through Learned Language Abstraction
- Safety Augmented Value Estimation from Demonstrations (SAVED): Safe Deep Model-Based RL for Sparse Cost Robotic Tasks
- Self-Imitation Advantage Learning
- Lifelong Robotic Reinforcement Learning by Retaining Experiences
- Universal Value Density Estimation for Imitation Learning and Goal-Conditioned Reinforcement Learning
- Learning a Multi-Modal Policy via Imitating Demonstrations with Mixed Behaviors
- Continuous Control with Deep Reinforcement Learning for Autonomous Vessels
- Expert-augmented actor-critic for ViZDoom and Montezumas Revenge
- TAAC: Temporally Abstract Actor-Critic for Continuous Control
- Language-Conditioned Goal Generation: a New Approach to Language Grounding for RL
- Hindsight Generative Adversarial Imitation Learning
- Uncertainty-sensitive Learning and Planning with Ensembles
- Attention-Privileged Reinforcement Learning
- Weakly-Supervised Reinforcement Learning for Controllable Behavior
- Generating Automatic Curricula via Self-Supervised Active Domain Randomization
- Developing a Simple Model for Sand-Tool Interaction and Autonomously Shaping Sand
- Generalizing Across Multi-Objective Reward Functions in Deep Reinforcement Learning
- Challenges of Context and Time in Reinforcement Learning: Introducing Space Fortress as a Benchmark
- Modularization of End-to-End Learning: Case Study in Arcade Games
- Meta Automatic Curriculum Learning
- How Much Do Unstated Problem Constraints Limit Deep Robotic Reinforcement Learning?
- Robust Multi-Modal Policies for Industrial Assembly via Reinforcement Learning and Demonstrations: A Large-Scale Study
- Learning to Manipulate Object Collections Using Grounded State Representations
- Solving Challenging Dexterous Manipulation Tasks With Trajectory Optimisation and Reinforcement Learning
- Learning Generalizable Locomotion Skills with Hierarchical Reinforcement Learning
- Deep Learning with Experience Ranking Convolutional Neural Network for Robot Manipulator
- Active Perception and Representation for Robotic Manipulation
- Adversarial Intrinsic Motivation for Reinforcement Learning
- Estimating Q(s,s') with Deep Deterministic Dynamics Gradients
- Predictive Coding for Boosting Deep Reinforcement Learning with Sparse Rewards
- Review, Analysis and Design of a Comprehensive Deep Reinforcement Learning Framework
- Conservative Data Sharing for Multi-Task Offline Reinforcement Learning
- Hierarchical Robot Navigation in Novel Environments using Rough 2-D Maps
- Learning to Solve a Rubik's Cube with a Dexterous Hand
- Proximal Policy Gradient: PPO with Policy Gradient
- Disentangling causal effects for hierarchical reinforcement learning
- Hierarchical reinforcement learning for efficient exploration and transfer
- Diversity-based Trajectory and Goal Selection with Hindsight Experience Replay
- ADER: Adaptively Distilled Exemplar Replay Towards Continual Learning for Session-based Recommendation
- Interactive Learning from Activity Description
- Variational Empowerment as Representation Learning for Goal-Based Reinforcement Learning
- Tutorial and Survey on Probabilistic Graphical Model and Variational Inference in Deep Reinforcement Learning
- Imaginary Hindsight Experience Replay: Curious Model-based Learning for Sparse Reward Tasks
- Forgetful Experience Replay in Hierarchical Reinforcement Learning from Demonstrations
- Hierarchical deep reinforcement learning controlled three-dimensional navigation of microrobots in blood vessels
- DeepSlicing: Deep Reinforcement Learning Assisted Resource Allocation for Network Slicing
- Unbiased Methods for Multi-Goal Reinforcement Learning
- Reinforcement Learning for Industrial Control Network Cyber Security Orchestration
- Quantum Compiling by Deep Reinforcement Learning
- Touch-based Curiosity for Sparse-Reward Tasks
- Mutual Information-based State-Control for Intrinsically Motivated Reinforcement Learning
- A Deep Learning Approach to Grasping the Invisible
- Efficient Robotic Task Generalization Using Deep Model Fusion Reinforcement Learning
- Multi-Robot Formation Control Using Reinforcement Learning
- Effects of sparse rewards of different magnitudes in the speed of learning of model-based actor critic methods
- Learning Adaptive Display Exposure for Real-Time Advertising
- Goal-oriented Trajectories for Efficient Exploration
- Accelerating Reinforcement Learning with a Directional-Gaussian-Smoothing Evolution Strategy
- Asynchronous Episodic Deep Deterministic Policy Gradient: Towards Continuous Control in Computationally Complex Environments
- Robotic self-representation improves manipulation skills and transfer learning
- Flexible and Efficient Long-Range Planning Through Curious Exploration
- Self-Paced Deep Reinforcement Learning
- Reinforcement Learning Control of Robotic Knee with Human in the Loop by Flexible Policy Iteration
- Evaluating a Generative Adversarial Framework for Information Retrieval
- Reinforcement Learning with Time-dependent Goals for Robotic Musicians
- Hindsight Experience Replay with Kronecker Product Approximate Curvature
- Deep Reinforcement Learning Based Robot Arm Manipulation with Efficient Training Data through Simulation
- Scaling All-Goals Updates in Reinforcement Learning Using Convolutional Neural Networks
- Action Redundancy in Reinforcement Learning
- Subgoal-based Reward Shaping to Improve Efficiency in Reinforcement Learning
- Self-supervised Reinforcement Learning with Independently Controllable Subgoals
- Hierarchical Reinforcement Learning Framework towards Multi-agent Navigation
- A survey of benchmarking frameworks for reinforcement learning
- Don't Do What Doesn't Matter: Intrinsic Motivation with Action Usefulness
- Value-Based Reinforcement Learning for Continuous Control Robotic Manipulation in Multi-Task Sparse Reward Settings
- SocialAI: Benchmarking Socio-Cognitive Abilities in Deep Reinforcement Learning Agents
- Auditing Robot Learning for Safety and Compliance during Deployment
- Using Logical Specifications of Objectives in Multi-Objective Reinforcement Learning
- Questions to Guide the Future of Artificial Intelligence Research
- Multi-task Reinforcement Learning with a Planning Quasi-Metric
- Deep Reinforcement Learning using Genetic Algorithm for Parameter Optimization
- Sim-to-Real Transfer of Robot Learning with Variable Length Inputs
- Evolutionary Stochastic Policy Distillation
- Automatic Goal Generation using Dynamical Distance Learning
- Direct Random Search for Fine Tuning of Deep Reinforcement Learning Policies
- When Autonomous Systems Meet Accuracy and Transferability through AI: A Survey
- Learn Proportional Derivative Controllable Latent Space from Pixels
- Contrastive Active Inference
- Generative Exploration and Exploitation
- Attribute-Based Robotic Grasping with One-Grasp Adaptation
- Sentiment Analysis for Reinforcement Learning
- Task-Oriented Language Grounding for Language Input with Multiple Sub-Goals of Non-Linear Order
- Unsupervised Skill-Discovery and Skill-Learning in Minecraft
- State Representation Learning from Demonstration
- Parallelized Reverse Curriculum Generation
- Reinforcement Learning for Robust Missile Autopilot Design
- Sample Efficiency in Sparse Reinforcement Learning: Or Your Money Back
- Biological Blueprints for Next Generation AI Systems
- ACDER: Augmented Curiosity-Driven Experience Replay
- Generalization in Text-based Games via Hierarchical Reinforcement Learning
- A Probabilistic Interpretation of Self-Paced Learning with Applications to Reinforcement Learning
- ViP: Video Platform for PyTorch
- Hierarchical Reinforcement Learning with Optimal Level Synchronization Based on Flow-Based Deep Generative Model
- Improving On-policy Learning with Statistical Reward Accumulation
- Locality-Sensitive Experience Replay for Online Recommendation