Agent Modelling under Partial Observability for Deep Reinforcement Learning
arXiv:2006.09447
Abstract
Modelling the behaviours of other agents is essential for understanding how agents interact and making effective decisions. Existing methods for agent modelling commonly assume knowledge of the local observations and chosen actions of the modelled agents during execution. To eliminate this assumption, we extract representations from the local information of the controlled agent using encoder-decoder architectures. Using the observations and actions of the modelled agents during training, our models learn to extract representations about the modelled agents conditioned only on the local observations of the controlled agent. The representations are used to augment the controlled agent's decision policy which is trained via deep reinforcement learning; thus, during execution, the policy does not require access to other agents' information. We provide a comprehensive evaluation and ablations studies in cooperative, competitive and mixed multi-agent environments, showing that our method achieves higher returns than baseline methods which do not use the learned representations.
Published in the 35th Conference on Neural Information Processing Systems (NeurIPS 2021)
References in corpus (15)
- Representation Learning with Contrastive Predictive Coding
- Learning a Generic Value-Selection Heuristic Inside a Constraint Programming Solver
- Recurrent World Models Facilitate Policy Evolution
- Autonomous Agents Modelling Other Agents: A Comprehensive Survey and Open Problems
- Learning to reinforcement learn
- Social Influence as Intrinsic Motivation for Multi-Agent Deep Reinforcement Learning
- Efficient Off-Policy Meta-Reinforcement Learning via Probabilistic Context Variables
- Dealing with Non-Stationarity in Multi-Agent Deep Reinforcement Learning
- Learning Plannable Representations with Causal InfoGAN
- Learning Invariant Representations for Reinforcement Learning without Reconstruction
- Learning Policy Representations in Multiagent Systems
- VariBAD: A Very Good Method for Bayes-Adaptive Deep RL via Meta-Learning
- Relational Forward Models for Multi-Agent Learning
- Efficient Model-Based Deep Reinforcement Learning with Variational State Tabulation
- Decoupling Dynamics and Reward for Transfer Learning