Towards Generalization and Simplicity in Continuous Control
arXiv:1703.02660
Abstract
This work shows that policies with simple linear and RBF parameterizations can be trained to solve a variety of continuous control tasks, including the OpenAI gym benchmarks. The performance of these trained policies are competitive with state of the art results, obtained with more elaborate parameterizations such as fully connected neural networks. Furthermore, existing training and testing scenarios are shown to be very limited and prone to over-fitting, thus giving rise to only trajectory-centric policies. Training with a diverse initial state distribution is shown to produce more global policies with better generalization. This allows for interactive control scenarios where the system recovers from large on-line perturbations; as shown in the supplementary video.
NIPS 2017, Project page: https://sites.google.com/view/simple-pol
References in corpus (1)
Cited by in corpus (7)
- Simple random search provides a competitive approach to reinforcement learning
- Reverse Curriculum Generation for Reinforcement Learning
- Discrete Sequential Prediction of Continuous Actions for Deep RL
- TD-Regularized Actor-Critic Methods
- Augmented Random Search for Quadcopter Control: An alternative to Reinforcement Learning
- Provable Regret Bounds for Deep Online Learning and Control
- Online Algorithms and Policies Using Adaptive and Machine Learning Approaches