Dynamics of Boltzmann Q-Learning in Two-Player Two-Action Games
arXiv:1109.1528 · doi:10.1103/PhysRevE.85.041145
Abstract
We consider the dynamics of Q-learning in two-player two-action games with a Boltzmann exploration mechanism. For any non-zero exploration rate the dynamics is dissipative, which guarantees that agent strategies converge to rest points that are generally different from the game's Nash Equlibria (NE). We provide a comprehensive characterization of the rest point structure for different games, and examine the sensitivity of this structure with respect to the noise due to exploration. Our results indicate that for a class of games with multiple NE the asymptotic behavior of learning dynamics can undergo drastic changes at critical exploration rates. Furthermore, we demonstrate that for certain games with a single NE, it is possible to have additional rest points (not corresponding to any NE) that persist for a finite range of the exploration rates and disappear when the exploration rates of both players tend to zero.
10 pages, 12 figures. Version 2: added more extensive discussion of asymmetric equilibria; clarified conditions for continuous/discontinuous bifurcations in coordination/anti-coordination games
Cited by in corpus (15)
- On the Properties of the Softmax Function with Application in Game Theory and Reinforcement Learning
- On Passivity, Reinforcement Learning and Higher-Order Learning in Multi-Agent Finite Games
- Cycle frequency in standard Rock-Paper-Scissors games: Evidence from experimental economics
- Artificial intelligence meets minority game: toward optimal resource allocation
- The dynamics of opinion expression
- Unlearnable Games and "Satisficing'' Decisions: A Simple Model for a Complex World
- Coevolutionary networks of reinforcement-learning agents
- Replicator dynamics with turnover of players
- Follow-the-Regularized-Leader Routes to Chaos in Routing Games
- On Improving Energy Efficiency within Green Femtocell Networks: A Hierarchical Reinforcement Learning Approach
- Coordination problems on networks revisited: statics and dynamics
- Improving Energy Efficiency in Femtocell Networks: A Hierarchical Reinforcement Learning Framework
- Unstable Dynamics of Adaptation in Unknown Environment due to Novelty Seeking
- Catastrophe by Design in Population Games: Destabilizing Wasteful Locked-in Technologies
- Adaptive Decision Making via Entropy Minimization