Cycles of cooperation and defection in imperfect learning
arXiv:1101.4378 · doi:10.1088/1742-5468/2011/08/P08007
Abstract
When people play a repeated game they usually try to anticipate their opponents' moves based on past observations, and then decide what action to take next. Behavioural economics studies the mechanisms by which strategic decisions are taken in these adaptive learning processes. We here investigate a model of learning the iterated prisoner's dilemma game. Players have the choice between three strategies, always defect (ALLD), always cooperate (ALLC) and tit-for-tat (TFT). The only strict Nash equilibrium in this situation is ALLD. When players learn to play this game convergence to the equilibrium is not guaranteed, for example we find cooperative behaviour if players discount observations in the distant past. When agents use small samples of observed moves to estimate their opponent's strategy the learning process is stochastic, and sustained oscillations between cooperation and defection can emerge. These cycles are similar to those found in stochastic evolutionary processes, but the origin of the noise sustaining the oscillations is different and lies in the imperfect sampling of the opponent's strategy. Based on a systematic expansion technique, we are able to predict the properties of these learning cycles, providing an analytical tool with which the outcome of more general stochastic adaptation processes can be characterised.
18 pages, 11 figures
References in corpus (11)
- Evolutionary games on graphs
- Mobility promotes and jeopardizes biodiversity in rock-paper-scissors games
- Coevolutionary Dynamics: From Finite to Infinite Populations
- Human strategy updating in evolutionary games
- Memory-Based Snowdrift Game on Networks
- Coexistence versus extinction in the stochastic cyclic Lotka-Volterra model
- Oscillatory Dynamics in Rock-Paper-Scissors Games with Mutations
- Effect of memory on the prisoner's dilemma game in a square lattice
- How limit cycles and quasi-cycles are related in systems with intrinsic noise
- Evolutionary dynamics, intrinsic noise and cycles of co-operation
- Intrinsic noise in game dynamical learning
Cited by in corpus (8)
- Deterministic limit of temporal difference reinforcement learning for stochastic games
- Limit Cycles Sparked by Mutation in the Repeated Prisoner's Dilemma
- Learning to play public good games
- Unlearnable Games and "Satisficing'' Decisions: A Simple Model for a Complex World
- Effects of noise on convergent game learning dynamics
- Coordination problems on networks revisited: statics and dynamics
- Fence-sitters Protect Cooperation in Complex Networks
- Unstable Dynamics of Adaptation in Unknown Environment due to Novelty Seeking