A Laplacian Framework for Option Discovery in Reinforcement Learning
arXiv:1703.00956
Abstract
Representation learning and option discovery are two of the biggest challenges in reinforcement learning (RL). Proto-value functions (PVFs) are a well-known approach for representation learning in MDPs. In this paper we address the option discovery problem by showing how PVFs implicitly define options. We do it by introducing eigenpurposes, intrinsic reward functions derived from the learned representations. The options discovered from eigenpurposes traverse the principal directions of the state space. They are useful for multiple tasks because they are discovered without taking the environment's rewards into consideration. Moreover, different options act at different time scales, making them helpful for exploration. We demonstrate features of eigenpurposes in traditional tabular domains as well as in Atari 2600 games.
Appearing in the Proceedings of the 34th International Conference on Machine Learning (ICML)
References in corpus (3)
Cited by in corpus (33)
- An information-theoretic perspective on intrinsic motivation in reinforcement learning: a survey
- Why Does Hierarchy (Sometimes) Work So Well in Reinforcement Learning?
- Composable Planning with Attributes
- A Survey on Recent Advances and Challenges in Reinforcement Learning Methods for Task-Oriented Dialogue Policy Learning
- A Geometric Perspective on Optimal Representations for Reinforcement Learning
- Induction and Exploitation of Subgoal Automata for Reinforcement Learning
- Discovering Options for Exploration by Minimizing Cover Time
- A Sufficient Statistic for Influence in Structured Multiagent Environments
- Finding Options that Minimize Planning Time
- Discovery of Options via Meta-Learned Subgoals
- Weakly-Supervised Reinforcement Learning for Controllable Behavior
- Hierarchical principles of embodied reinforcement learning: A review
- A Survey of Exploration Methods in Reinforcement Learning
- IR-VIC: Unsupervised Discovery of Sub-goals for Transfer in RL
- Efficient Exploration through Intrinsic Motivation Learning for Unsupervised Subgoal Discovery in Model-Free Hierarchical Reinforcement Learning
- Temporally-Extended ε-Greedy Exploration
- On The Effect of Auxiliary Tasks on Representation Dynamics
- LISPR: An Options Framework for Policy Reuse with Reinforcement Learning
- DROP: Deep relocating option policy for optimal ride-hailing vehicle repositioning
- Memory-Efficient Episodic Control Reinforcement Learning with Dynamic Online k-means
- Hierarchical Representation Learning for Markov Decision Processes
- Abstract Value Iteration for Hierarchical Reinforcement Learning
- Sample-Efficient Reinforcement Learning with Maximum Entropy Mellowmax Episodic Control
- Temporal Abstraction in Reinforcement Learning with the Successor Representation
- Flexible and Efficient Long-Range Planning Through Curious Exploration
- Hierarchical Skills for Efficient Exploration
- Provable Hierarchy-Based Meta-Reinforcement Learning
- Average-Reward Learning and Planning with Options
- Direct then Diffuse: Incremental Unsupervised Skill Discovery for State Covering and Goal Reaching
- Playing Atari Ball Games with Hierarchical Reinforcement Learning
- Variational Intrinsic Control Revisited
- Augmenting the action space with conventions to improve multi-agent cooperation in Hanabi
- Option Discovery in the Absence of Rewards with Manifold Analysis