Empathy Modeling in Active Inference Agents for Perspective-Taking and Alignment
arXiv:2602.20936
Abstract
Artificial agents that model other agents must predict their behavior and determine whether their outcomes matter within action selection. We introduce an active inference framework that separates these components by combining a history-conditioned Theory of Mind model with an explicit other-regarding valuation parameter, . We instantiate the framework in the Iterated Prisoner's Dilemma. The joint empathy configuration reorganizes the long-run cooperation landscape: sufficiently strong and symmetric other-regarding valuation supports sustained mutual cooperation, whereas strong asymmetry exposes the more empathic agent to systematic exploitation. Along the symmetric diagonal, cooperation exhibits a sharp but continuous finite-precision crossover. Fixed-partner sweeps reveal that the apparent cooperation boundary is path-dependent and that temporal variability is elevated where those paths cross it. Online Bayesian inference over opponent parameters modestly facilitates cooperation near the behavioral boundary but does not substitute for other-regarding valuation. Direct model comparison likewise shows that opponent-sensitive prediction at does not generate cooperation. Planning depth has a partner-dependent effect: it slightly reduces cooperation when modeled reciprocity is weak but strongly increases cooperation against a reciprocating partner such as tit-for-tat. These results distinguish prediction, planning, and prosocial valuation as separable but interacting components of social agency. They also reveal a central limitation of unconditional empathic concern: the same valuation that stabilizes mutual cooperation creates predictable vulnerability when concern is not reciprocated.
Code and data: https://doi.org/10.5281/zenodo.21908008