Revisiting Design Choices in Proximal Policy Optimization
arXiv:2009.10897
Abstract
Proximal Policy Optimization (PPO) is a popular deep policy gradient algorithm. In standard implementations, PPO regularizes policy updates with clipped probability ratios, and parameterizes policies with either continuous Gaussian distributions or discrete Softmax distributions. These design choices are widely accepted, and motivated by empirical performance comparisons on MuJoCo and Atari benchmarks. We revisit these practices outside the regime of current benchmarks, and expose three failure modes of standard PPO. We explain why standard design choices are problematic in these cases, and show that alternative choices of surrogate objectives and policy parameterizations can prevent the failure modes. We hope that our work serves as a reminder that many algorithmic design choices in reinforcement learning are tied to specific simulation environments. We should not implicitly accept these choices as a standard part of a more general algorithm.
References in corpus (15)
- Trust Region Policy Optimization
- Dota 2 with Large Scale Deep Reinforcement Learning
- Solving Rubik's Cube with a Robot Hand
- Leveraging Demonstrations for Deep Reinforcement Learning on Robotics Problems with Sparse Rewards
- Parameter Space Noise for Exploration
- Deep Reinforcement Learning that Matters
- #Exploration: A Study of Count-Based Exploration for Deep Reinforcement Learning
- Learning by Playing - Solving Sparse Reward Tasks from Scratch
- Deep Exploration via Randomized Value Functions
- Chip Placement with Deep Reinforcement Learning
- On the Theory of Policy Gradient Methods: Optimality, Approximation, and Distribution Shift
- What Matters In On-Policy Reinforcement Learning? A Large-Scale Empirical Study
- A unified view of entropy-regularized Markov decision processes
- Neural Proximal/Trust Region Policy Optimization Attains Globally Optimal Policy
- V-MPO: On-Policy Maximum a Posteriori Policy Optimization for Discrete and Continuous Control