What Matters In On-Policy Reinforcement Learning? A Large-Scale Empirical Study
arXiv:2006.05990
Abstract
In recent years, on-policy reinforcement learning (RL) has been successfully applied to many different continuous control tasks. While RL algorithms are often conceptually simple, their state-of-the-art implementations take numerous low- and high-level design decisions that strongly affect the performance of the resulting agents. Those choices are usually not extensively discussed in the literature, leading to discrepancy between published descriptions of algorithms and their implementations. This makes it hard to attribute progress in RL and slows down overall progress [Engstrom'20]. As a step towards filling that gap, we implement >50 such ``choices'' in a unified on-policy RL framework, allowing us to investigate their impact in a large-scale empirical study. We train over 250'000 agents in five continuous control environments of different complexity and provide insights and practical recommendations for on-policy training of RL agents.
References in corpus (17)
- Adam: A Method for Stochastic Optimization
- Playing Atari with Deep Reinforcement Learning
- Soft Actor-Critic: Off-Policy Maximum Entropy Deep Reinforcement Learning with a Stochastic Actor
- Soft Actor-Critic Algorithms and Applications
- High-Dimensional Continuous Control Using Generalized Advantage Estimation
- Scaling Laws for Neural Language Models
- Challenging Common Assumptions in the Unsupervised Learning of Disentangled Representations
- Solving Rubik's Cube with a Robot Hand
- Learning Dexterous In-Hand Manipulation
- Maximum a Posteriori Policy Optimisation
- Reproducibility of Benchmarked Deep Reinforcement Learning Tasks for Continuous Control
- Advantage-Weighted Regression: Simple and Scalable Off-Policy Reinforcement Learning
- When to Trust Your Model: Model-Based Policy Optimization
- The Mirage of Action-Dependent Baselines in Reinforcement Learning
- V-MPO: On-Policy Maximum a Posteriori Policy Optimization for Discrete and Continuous Control
- SEED RL: Scalable and Efficient Deep-RL with Accelerated Central Inference
- Time Limits in Reinforcement Learning
Cited by in corpus (23)
- Autonomous Unmanned Aerial Vehicle Navigation using Reinforcement Learning: A Systematic Review
- What Matters in Learning from Offline Human Demonstrations for Robot Manipulation
- Phasic Policy Gradient
- Learning to Locomote: Understanding How Environment Design Matters for Deep Reinforcement Learning
- Rethinking the Implementation Tricks and Monotonicity Constraint in Cooperative Multi-Agent Reinforcement Learning
- Applications of Generative AI in Healthcare: algorithmic, ethical, legal and societal considerations
- Benchmarking Potential Based Rewards for Learning Humanoid Locomotion
- What Matters for Adversarial Imitation Learning?
- Revisiting Design Choices in Proximal Policy Optimization
- Observation Space Matters: Benchmark and Optimization Algorithm
- Zero-Shot Terrain Generalization for Visual Locomotion Policies
- Differentiable Trust Region Layers for Deep Reinforcement Learning
- A Deeper Look at Discounting Mismatch in Actor-Critic Algorithms
- First-Order Problem Solving through Neural MCTS based Reinforcement Learning
- Reward Function Design for Crowd Simulation via Reinforcement Learning
- A Pragmatic Look at Deep Imitation Learning
- A Continuous Optimisation Benchmark Suite from Neural Network Regression
- MimicBot: Combining Imitation and Reinforcement Learning to win in Bot Bowl
- Faster Policy Learning with Continuous-Time Gradients
- Average-Reward Reinforcement Learning with Trust Region Methods
- Learning Representations for Pixel-based Control: What Matters and Why?
- Comparing Reinforcement Learning and Human Learning using the Game of Hidden Rules
- Correcting Momentum in Temporal Difference Learning