A Survey of Learning in Multiagent Environments: Dealing with Non-Stationarity
arXiv:1707.09183
Abstract
The key challenge in multiagent learning is learning a best response to the behaviour of other agents, which may be non-stationary: if the other agents adapt their strategy as well, the learning target moves. Disparate streams of research have approached non-stationarity from several angles, which make a variety of implicit assumptions that make it hard to keep an overview of the state of the art and to validate the innovation and significance of new works. This survey presents a coherent overview of work that addresses opponent-induced non-stationarity with tools from game theory, reinforcement learning and multi-armed bandits. Further, we reflect on the principle approaches how algorithms model and cope with this non-stationarity, arriving at a new framework and five categories (in increasing order of sophistication): ignore, forget, respond to target models, learn models, and theory of mind. A wide range of state-of-the-art algorithms is classified into a taxonomy, using these categories and key characteristics of the environment (e.g., observability) and adaptation behaviour of the opponents (e.g., smooth, abrupt). To clarify even further we present illustrative variations of one domain, contrasting the strengths and limitations of each category. Finally, we discuss in which environments the different approaches yield most merit, and point to promising avenues of future research.
64 pages, 7 figures. Under review since November 2016
References in corpus (7)
- Deep Reinforcement Learning for Multi-Agent Systems: A Review of Challenges, Solutions and Applications
- A Survey and Critique of Multiagent Deep Reinforcement Learning
- Stabilising Experience Replay for Deep Multi-Agent Reinforcement Learning
- Multi-agent Reinforcement Learning in Sequential Social Dilemmas
- Opponent Modeling in Deep Reinforcement Learning
- A Polynomial-time Nash Equilibrium Algorithm for Repeated Stochastic Games
- Automated Planning in Repeated Adversarial Games
Cited by in corpus (48)
- Deep Reinforcement Learning for Multi-Agent Systems: A Review of Challenges, Solutions and Applications
- A Survey and Critique of Multiagent Deep Reinforcement Learning
- Autonomous Agents Modelling Other Agents: A Comprehensive Survey and Open Problems
- Game-Theoretic Multiagent Reinforcement Learning
- Multi-Objective Multi-Agent Decision Making: A Utility-based Analysis and Survey
- Dealing with Non-Stationarity in Multi-Agent Deep Reinforcement Learning
- R-MADDPG for Partially Observable Environments and Limited Communication
- A Deep Policy Inference Q-Network for Multi-Agent Systems
- Off-Policy Multi-Agent Decomposed Policy Gradients
- Decentralized Learning for Optimality in Stochastic Dynamic Teams and Games with Local Control and Global State Information
- Deep Reinforcement Learning in Computer Vision: A Comprehensive Survey
- Delay-Aware Multi-Agent Reinforcement Learning for Cooperative and Competitive Environments
- Quantifying the Impact of Non-Stationarity in Reinforcement Learning-Based Traffic Signal Control
- Toward Packet Routing with Fully-distributed Multi-agent Deep Reinforcement Learning
- Applications of Multi-Agent Reinforcement Learning in Future Internet: A Comprehensive Survey
- On Multi-Agent Learning in Team Sports Games
- Algorithms in Multi-Agent Systems: A Holistic Perspective from Reinforcement Learning and Game Theory
- A Policy Gradient Algorithm for Learning to Learn in Multiagent Reinforcement Learning
- Learning Multi-agent Communication under Limited-bandwidth Restriction for Internet Packet Routing
- Winning Isn't Everything: Enhancing Game Development with Intelligent Agents
- Learning Agent Communication under Limited Bandwidth by Message Pruning
- Emergence of Theory of Mind Collaboration in Multiagent Systems
- Learning in Nonzero-Sum Stochastic Games with Potentials
- Revisiting Parameter Sharing in Multi-Agent Deep Reinforcement Learning
- Decentralized Q-Learning in Zero-sum Markov Games
- Provable Fictitious Play for General Mean-Field Games
- Deception in Social Learning: A Multi-Agent Reinforcement Learning Perspective
- Emergent Road Rules In Multi-Agent Driving Environments
- Influencing Towards Stable Multi-Agent Interactions
- Incorporating Pragmatic Reasoning Communication into Emergent Language
- Coordination-driven learning in multi-agent problem spaces
- PsiPhi-Learning: Reinforcement Learning with Demonstrations using Successor Features and Inverse Temporal Difference Learning
- Event-Triggered Multi-agent Reinforcement Learning with Communication under Limited-bandwidth Constraint
- Agent Modelling under Partial Observability for Deep Reinforcement Learning
- Multi-user Resource Control with Deep Reinforcement Learning in IoT Edge Computing
- Regret Bounds for Decentralized Learning in Cooperative Multi-Agent Dynamical Systems
- A Game-Theoretic Approach to Multi-Agent Trust Region Optimization
- Towards Efficient Detection and Optimal Response against Sophisticated Opponents
- The Evolutionary Dynamics of Independent Learning Agents in Population Games
- Multi-Agent Deep Reinforcement Learning with Adaptive Policies
- Health-Informed Policy Gradients for Multi-Agent Reinforcement Learning
- Decentralized Reinforcement Learning for Multi-Target Search and Detection by a Team of Drones
- Detecting Rewards Deterioration in Episodic Reinforcement Learning
- Coordinated Proximal Policy Optimization
- Interactive AI with a Theory of Mind
- Behaviour-conditioned policies for cooperative reinforcement learning tasks
- The Confluence of Networks, Games and Learning
- The Power of Communication in a Distributed Multi-Agent System