paper

Advanced Policies: A First-Principles Path from Policy Gradient to Q-Learning

arXiv:2005.08844

Abstract

We apply entropy to reinforcement learning in two ways: entropy-augmented reward defines the soft objective, while relative-entropy regularization controls the size of a policy-improvement step without changing that objective. Together they yield a one-parameter family of advanced policies connecting a base policy to its soft-greedy policy. Under exact evaluation, every nontrivial member improves upon the common base, although the improvement need not be monotone along the family. Under policy-value consistency, this family forms an entropic mirror-descent path. We then relax this consistency, treating the policy and action-value function as independent coordinates of the advanced policy. Differentiating the objective of the advanced policy yields Advanced Actor-Critic (AAC), whose endpoint gradients recover soft policy gradient and an action-centered Q-learning-like update. We develop implementations for discrete and continuous actions. Experiments demonstrate effective learning at both endpoints and intermediate parameter values.

Substantially revised version. Corrects the earlier monotonic-improvement claim, develops a substantially revised formulation of Advanced Actor-Critic (AAC), and adds expanded theoretical analysis, continuous-action implementations, and substantially expanded experiments

References in corpus (5)