paper

Reinforcement learning with reputation-based adaptive exploration promotes cooperation

arXiv:2604.08103 · doi:10.1063/5.0348706

Abstract

Reinforcement learning provides a framework for studying how individuals adjust their behavior through repeated interaction and feedback in social dilemmas. In Q-learning, exploration controls how often agents choose actions other than those favored by their current learned Q-values. Yet existing models usually treat the exploration rate as a constant parameter. In systems with social evaluation, however, trial-and-error behavior carries different costs and opportunities for agents with different reputations, making exploration dependent on social standing rather than uniform across agents. Herein, we develop a spatial prisoner's dilemma model in which Q-learning agents adapt their exploration rates according to local reputation differences, while reputation is updated through an asymmetric, state-dependent rule. Results show that adaptive exploration and asymmetric reputation updating each promote cooperation, but their combination produces a stronger reinforcing effect than either mechanism alone. Low-reputation agents explore more and can recover reputation through cooperation, while high-reputation agents explore less and avoid reputation losses caused by defection. This mechanism also reorganizes cooperation in space, producing a stable checkerboard-like coexistence at intermediate reputation concern. In addition, cooperation is most vulnerable at intermediate baseline exploration rates, whereas stronger asymmetric reputation updating mitigates this exploration-induced disruption. These results suggest that reputation can act not only as a record of past behavior, but also as a dynamic signal that regulates exploratory behavior during learning and thereby stabilizes cooperation.

12 pages, 6 figures

Reinforcement learning with reputation-based adaptive exploration promotes cooperation · wovepaper