Multi-Advisor Reinforcement Learning
arXiv:1704.00756
Abstract
We consider tackling a single-agent RL problem by distributing it to learners. These learners, called advisors, endeavour to solve the problem from a different focus. Their advice, taking the form of action values, is then communicated to an aggregator, which is in control of the system. We show that the local planning method for the advisors is critical and that none of the ones found in the literature is flawless: the egocentric planning overestimates values of states where the other advisors disagree, and the agnostic planning is inefficient around danger zones. We introduce a novel approach called empathic and discuss its theoretical aspects. We empirically examine and validate our theoretical findings on a fruit collection task.
Submitted at ICLR2018
References in corpus (4)
Cited by in corpus (6)
- Hybrid Reward Architecture for Reinforcement Learning
- A Meta-MDP Approach to Exploration for Lifelong Reinforcement Learning
- A Deeper Look at Discounting Mismatch in Actor-Critic Algorithms
- Danger-aware Adaptive Composition of DRL Agents for Self-navigation
- On mechanisms for transfer using landmark value functions in multi-task lifelong reinforcement learning
- Prioritized Soft Q-Decomposition for Lexicographic Reinforcement Learning