Reinforcement Learning Produces Dominant Strategies for the Iterated Prisoner's Dilemma
arXiv:1707.06307 · doi:10.1371/journal.pone.0188046
Abstract
We present tournament results and several powerful strategies for the Iterated Prisoner's Dilemma created using reinforcement learning techniques (evolutionary and particle swarm algorithms). These strategies are trained to perform well against a corpus of over 170 distinct opponents, including many well-known and classic strategies. All the trained strategies win standard tournaments against the total collection of other opponents. The trained strategies and one particular human made designed strategy are the top performers in noisy tournaments also.
References in corpus (5)
- The NumPy array: a structure for efficient numerical computation
- Evolutionary Algorithms for Reinforcement Learning
- Reinforcement Learning Produces Dominant Strategies for the Iterated Prisoner's Dilemma
- An open reproducible framework for the study of the iterated prisoner's dilemma
- On some winning strategies for the Iterated Prisoner's Dilemma or Mr. Nice Guy and the Cosa Nostra
Cited by in corpus (14)
- Social physics
- Reinforcement Learning Produces Dominant Strategies for the Iterated Prisoner's Dilemma
- Symmetric equilibrium of multi-agent reinforcement learning in repeated prisoner's dilemma
- Emergence of Cooperation in Two-agent Repeated Games with Reinforcement Learning
- Evolution Reinforces Cooperation with the Emergence of Self-Recognition Mechanisms: an empirical study of the Moran process for the iterated Prisoner's dilemma
- Modeling Moral Choices in Social Dilemmas with Multi-Agent Reinforcement Learning
- Memory-two strategies forming symmetric mutual reinforcement learning equilibrium in repeated prisoners' dilemma game
- Learning in two-player games between transparent opponents
- Preference-based opponent shaping in differentiable games
- Memory depth of finite state machine strategies for the iterated prisoner's dilemma
- Learning multiagent coordination in the absence of communication channels
- Top Score in Axelrod Tournament
- A Predictive Strategy for the Iterated Prisoner's Dilemma
- Emergence and Stability of Self-Evolved Cooperative Strategies using Stochastic Machines