Advanced Policies: A First-Principles Path from Policy Gradient to Q-Learning
arXiv:2005.08844
Abstract
We apply entropy to reinforcement learning in two ways: entropy-augmented reward defines the soft objective, while relative-entropy regularization controls the size of a policy-improvement step without changing that objective. Together they yield a one-parameter family of advanced policies connecting a base policy to its soft-greedy policy. Under exact evaluation, every nontrivial member improves upon the common base, although the improvement need not be monotone along the family. Under policy-value consistency, this family forms an entropic mirror-descent path. We then relax this consistency, treating the policy and action-value function as independent coordinates of the advanced policy. Differentiating the objective of the advanced policy yields Advanced Actor-Critic (AAC), whose endpoint gradients recover soft policy gradient and an action-centered Q-learning-like update. We develop implementations for discrete and continuous actions. Experiments demonstrate effective learning at both endpoints and intermediate parameter values.
Substantially revised version. Corrects the earlier monotonic-improvement claim, develops a substantially revised formulation of Advanced Actor-Critic (AAC), and adds expanded theoretical analysis, continuous-action implementations, and substantially expanded experiments
References in corpus (5)
- Soft Actor-Critic: Off-Policy Maximum Entropy Deep Reinforcement Learning with a Stochastic Actor
- Addressing Function Approximation Error in Actor-Critic Methods
- Reinforcement Learning with Deep Energy-Based Policies
- Bridging the Gap Between Value and Policy Based Reinforcement Learning
- Equivalence Between Policy Gradients and Soft Q-Learning