artificial intelligence

Calculating Mutual Information between a Reward Maximizer and its Environment

arXiv:2602.12963

summary

The paper derives an exact information-theoretic bound showing that an optimal deterministic policy in a Controlled Markov Process reveals n log m bits about the environment, quantifying the implicit world model required for optimality.

Abstract

An important question in the field of AI is the extent to which successful behaviour requires an internal representation of the world. In this work, we quantify the amount of information an optimal policy provides about the underlying environment. We consider a Controlled Markov Process (CMP) with states and actions, assuming a uniform prior over the space of possible transition dynamics. We prove that observing a deterministic policy that is optimal for any non-constant reward function then conveys exactly bits of information about the environment. Specifically, we show that the mutual information between the environment and the optimal policy is bits. This bound holds across a broad class of objectives, including finite-horizon, infinite-horizon discounted, and time-averaged reward maximization. These findings provide a precise information-theoretic lower bound on the ``implicit world model'' necessary for optimality.

Topics & keywords

#reinforcement learning#information theory#optimal policy#controlled markov process#mutual informationmutual informationoptimal policycontrolled Markov processuniform priorn log m bitsreward maximization