Maximum Entropy-Regularized Multi-Goal Reinforcement Learning
arXiv:1905.08786
Abstract
In Multi-Goal Reinforcement Learning, an agent learns to achieve multiple goals with a goal-conditioned policy. During learning, the agent first collects the trajectories into a replay buffer, and later these trajectories are selected randomly for replay. However, the achieved goals in the replay buffer are often biased towards the behavior policies. From a Bayesian perspective, when there is no prior knowledge about the target goal distribution, the agent should learn uniformly from diverse achieved goals. Therefore, we first propose a novel multi-goal RL objective based on weighted entropy. This objective encourages the agent to maximize the expected return, as well as to achieve more diverse goals. Secondly, we developed a maximum entropy-based prioritization framework to optimize the proposed objective. For evaluation of this framework, we combine it with Deep Deterministic Policy Gradient, both with or without Hindsight Experience Replay. On a set of multi-goal robotic tasks of OpenAI Gym, we compare our method with other baselines and show promising improvements in both performance and sample-efficiency.
Published in International Conference on Machine Learning (ICML 2019), Long Beach, USA. arXiv admin note: text overlap with arXiv:1902.08039
Cited by in corpus (18)
- Exploration in Deep Reinforcement Learning: From Single-Agent to Multiagent Domain
- Skew-Fit: State-Covering Self-Supervised Reinforcement Learning
- An information-theoretic perspective on intrinsic motivation in reinforcement learning: a survey
- Curiosity-Driven Experience Prioritization via Density Estimation
- Variational Dynamic for Self-Supervised Exploration in Deep Reinforcement Learning
- World Model as a Graph: Learning Latent Landmarks for Planning
- Hindsight Goal Ranking on Replay Buffer for Sparse Reward Environment
- MHER: Model-based Hindsight Experience Replay
- C-Learning: Learning to Achieve Goals via Recursive Classification
- Entropy-Aware Model Initialization for Effective Exploration in Deep Reinforcement Learning
- Mutual Information State Intrinsic Control
- Tutorial and Survey on Probabilistic Graphical Model and Variational Inference in Deep Reinforcement Learning
- Exploration via Hindsight Goal Generation
- Mutual Information-based State-Control for Intrinsically Motivated Reinforcement Learning
- Density-based Curriculum for Multi-goal Reinforcement Learning with Sparse Rewards
- Bias-reduced Multi-step Hindsight Experience Replay for Efficient Multi-goal Reinforcement Learning
- Planning under Uncertainty to Goal Distributions
- Adaptive Multi-Goal Exploration