Adversarially Guided Actor-Critic
arXiv:2102.04376
Abstract
Despite definite success in deep reinforcement learning problems, actor-critic algorithms are still confronted with sample inefficiency in complex environments, particularly in tasks where efficient exploration is a bottleneck. These methods consider a policy (the actor) and a value function (the critic) whose respective losses are built using different motivations and approaches. This paper introduces a third protagonist: the adversary. While the adversary mimics the actor by minimizing the KL-divergence between their respective action distributions, the actor, in addition to learning to solve the task, tries to differentiate itself from the adversary predictions. This novel objective stimulates the actor to follow strategies that could not have been correctly predicted from previous trajectories, making its behavior innovative in tasks where the reward is extremely rare. Our experimental analysis shows that the resulting Adversarially Guided Actor-Critic (AGAC) algorithm leads to more exhaustive exploration. Notably, AGAC outperforms current state-of-the-art methods on a set of various hard-exploration and procedurally-generated tasks.
Accepted at ICLR 2021
References in corpus (16)
- Trust Region Policy Optimization
- Scalable trust-region method for deep reinforcement learning using Kronecker-factored approximation
- A Study on Overfitting in Deep Reinforcement Learning
- Go-Explore: a New Approach for Hard-Exploration Problems
- Illuminating Generalization in Deep Reinforcement Learning through Procedural Level Generation
- Generalization and Regularization in DQN
- Connecting Generative Adversarial Networks and Actor-Critic Methods
- A Theory of Regularized Markov Decision Processes
- A Dissection of Overfitting and Generalization in Continuous Reinforcement Learning
- Learning to Understand Goal Specifications by Modelling Reward
- Generalization in Reinforcement Learning with Selective Noise Injection and Information Bottleneck
- Munchausen Reinforcement Learning
- The Reactor: A fast and sample-efficient Actor-Critic agent for Reinforcement Learning
- Self-Attentional Credit Assignment for Transfer in Reinforcement Learning
- Diversity Actor-Critic: Sample-Aware Entropy Regularization for Sample-Efficient Exploration
- Only Relevant Information Matters: Filtering Out Noisy Samples to Boost RL