Emergence of cooperation under punishment: A reinforcement learning perspective
arXiv:2401.16073 · doi:10.1063/5.0215702
Abstract
Punishment is a common tactic to sustain cooperation and has been extensively studied for a long time. While most of previous game-theoretic work adopt the imitation learning where players imitate the strategies who are better off, the learning logic in the real world is often much more complex. In this work, we turn to the reinforcement learning paradigm, where individuals make their decisions based upon their past experience and long-term returns. Specifically, we investigate the Prisoners' dilemma game with Q-learning algorithm, and cooperators probabilistically pose punishment on defectors in their neighborhood. Interestingly, we find that punishment could lead to either continuous or discontinuous cooperation phase transitions, and the nucleation process of cooperation clusters is reminiscent of the liquid-gas transition. The uncovered first-order phase transition indicates that great care needs to be taken when implementing the punishment compared to the continuous scenario.
7 pages, 6 figures
References in corpus (19)
- Statistical physics of human cooperation
- Social diversity and promotion of cooperation in the spatial prisoner's dilemma game
- Reward and cooperation in the spatial public goods game
- Phase diagrams for the spatial public goods game with pool-punishment
- Punish, but not too hard: How costly punishment spreads in the spatial public goods game
- Evolutionary establishment of moral and double moral standards through spatial interactions
- Memory-Based Snowdrift Game on Networks
- Probabilistic sharing solves the problem of costly punishment
- Promoting cooperation in social dilemmas via simple coevolutionary rules
- Restricted connections among distinguished players support cooperation
- Effectiveness of conditional punishment for the evolution of public cooperation
- Defector-accelerated cooperativeness and punishment in public goods games with mutations
- Effect of memory on the prisoner's dilemma game in a square lattice
- Dynamically generated cyclic dominance in spatial prisoner's dilemma games
- Voluntary rewards mediate the evolution of pool punishment for maintaining public goods in large populations
- Emergence of Cooperation in Two-agent Repeated Games with Reinforcement Learning
- Decoding trust: A reinforcement learning perspective
- Emergence of cooperation in a population with bimodal response behaviors
- Oscillatory cooperation prevalence emerges from misperception