paper

Logit Dynamics in Softmax Policy Gradient Methods

arXiv:2506.12912

Abstract

We analyzes the logit dynamics of softmax policy gradient methods. We derive the exact formula for the L2 norm of the logit update vector: This equation demonstrates that update magnitudes are determined by the chosen action's probability () and the policy's collision probability (), a measure of concentration inversely related to entropy. Our analysis reveals an inherent self-regulation mechanism where learning vigor is automatically modulated by policy confidence, providing a foundational insight into the stability and convergence of these methods.

7 pages