Reproducibility of Benchmarked Deep Reinforcement Learning Tasks for Continuous Control
arXiv:1708.04133
Abstract
Policy gradient methods in reinforcement learning have become increasingly prevalent for state-of-the-art performance in continuous control tasks. Novel methods typically benchmark against a few key algorithms such as deep deterministic policy gradients and trust region policy optimization. As such, it is important to present and use consistent baselines experiments. However, this can be difficult due to general variance in the algorithms, hyper-parameter tuning, and environment stochasticity. We investigate and discuss: the significance of hyper-parameters in policy gradients for continuous control, general variance in the algorithms, and reproducibility of reported results. We provide guidelines on reporting novel results as comparisons against baseline methods such that future researchers can make informed decisions when investigating novel methods.
Accepted to Reproducibility in Machine Learning Workshop, ICML'17
References in corpus (3)
Cited by in corpus (13)
- DeepMind Control Suite
- Benchmarking Model-Based Reinforcement Learning
- Boosting Soft Actor-Critic: Emphasizing Recent Experience without Forgetting the Past
- On the Global Convergence of Actor-Critic: A Case for Linear Quadratic Regulator with Ergodic Cost
- Q-Learning for Continuous Actions with Cross-Entropy Guided Policies
- DropoutDAgger: A Bayesian Approach to Safe Imitation Learning
- Map-based Multi-Policy Reinforcement Learning: Enhancing Adaptability of Robots by Deep Reinforcement Learning
- SLM Lab: A Comprehensive Benchmark and Modular Software Framework for Reproducible Deep Reinforcement Learning
- Catalyst.RL: A Distributed Framework for Reproducible RL Research
- Recruitment-imitation Mechanism for Evolutionary Reinforcement Learning
- Randomized Adversarial Imitation Learning for Autonomous Driving
- D2C 2.0: Decoupled Data-Based Approach for Learning to Control Stochastic Nonlinear Systems via Model-Free ILQR
- Doubly Robust Off-Policy Actor-Critic Algorithms for Reinforcement Learning