Return-Based Contrastive Representation Learning for Reinforcement Learning
arXiv:2102.10960
Abstract
Recently, various auxiliary tasks have been proposed to accelerate representation learning and improve sample efficiency in deep reinforcement learning (RL). However, existing auxiliary tasks do not take the characteristics of RL problems into consideration and are unsupervised. By leveraging returns, the most important feedback signals in RL, we propose a novel auxiliary task that forces the learnt representations to discriminate state-action pairs with different returns. Our auxiliary loss is theoretically justified to learn representations that capture the structure of a new form of state-action abstraction, under which state-action pairs with similar return distributions are aggregated together. In low data regime, our algorithm outperforms strong baselines on complex tasks in Atari games and DeepMind Control suite, and achieves even better performance when combined with existing auxiliary tasks.
ICLR 2021
References in corpus (7)
- Learning to Navigate in Complex Environments
- Loss is its own Reward: Self-Supervision for Reinforcement Learning
- DeepMDP: Learning Continuous Latent Space Models for Representation Learning
- Bootstrap Latent-Predictive Representations for Multitask Reinforcement Learning
- Universal Successor Features Approximators
- Auxiliary-task Based Deep Reinforcement Learning for Participant Selection Problem in Mobile Crowdsourcing
- Kinematic State Abstraction and Provably Efficient Rich-Observation Reinforcement Learning