Curiosity-Driven Experience Prioritization via Density Estimation
arXiv:1902.08039
Abstract
In Reinforcement Learning (RL), an agent explores the environment and collects trajectories into the memory buffer for later learning. However, the collected trajectories can easily be imbalanced with respect to the achieved goal states. The problem of learning from imbalanced data is a well-known problem in supervised learning, but has not yet been thoroughly researched in RL. To address this problem, we propose a novel Curiosity-Driven Prioritization (CDP) framework to encourage the agent to over-sample those trajectories that have rare achieved goal states. The CDP framework mimics the human learning process and focuses more on relatively uncommon events. We evaluate our methods using the robotic environment provided by OpenAI Gym. The environment contains six robot manipulation tasks. In our experiments, we combined CDP with Deep Deterministic Policy Gradient (DDPG) with or without Hindsight Experience Replay (HER). The experimental results show that CDP improves both performance and sample-efficiency of reinforcement learning agents, compared to state-of-the-art methods.
Accepted by NIPS Deep RL Workshop, 2018, link: https://sites.google.com/view/deep-rl-workshop-nips-2018 . arXiv admin note: substantial text overlap with arXiv:1810.01363 and text overlap with arXiv:1905.08786
References in corpus (1)
Cited by in corpus (20)
- Exploration in Deep Reinforcement Learning: A Survey
- A survey on intrinsic motivation in reinforcement learning
- Skew-Fit: State-Covering Self-Supervised Reinforcement Learning
- Boosting Soft Actor-Critic: Emphasizing Recent Experience without Forgetting the Past
- Perfect density models cannot guarantee anomaly detection
- Sample-efficient Reinforcement Learning Representation Learning with Curiosity Contrastive Forward Dynamics Model
- Striving for Simplicity and Performance in Off-Policy DRL: Output Normalization and Non-Uniform Sampling
- CCLF: A Contrastive-Curiosity-Driven Learning Framework for Sample-Efficient Reinforcement Learning
- Soft Hindsight Experience Replay
- Curiosity-Driven Multi-Criteria Hindsight Experience Replay
- Hindsight Generative Adversarial Imitation Learning
- Mutual Information-based State-Control for Intrinsically Motivated Reinforcement Learning
- Complex Robotic Manipulation via Graph-Based Hindsight Goal Generation
- Improved Exploring Starts by Kernel Density Estimation-Based State-Space Coverage Acceleration in Reinforcement Learning
- GAN-based Intrinsic Exploration For Sample Efficient Reinforcement Learning
- Off-policy Reinforcement Learning with Optimistic Exploration and Distribution Correction
- Adaptive Experience Selection for Policy Gradient
- Fixed -VAE Encoding for Curious Exploration in Complex 3D Environments
- Knowledge is reward: Learning optimal exploration by predictive reward cashing
- Exploring More When It Needs in Deep Reinforcement Learning