Discrete Sequential Prediction of Continuous Actions for Deep RL
arXiv:1705.05035
Abstract
It has long been assumed that high dimensional continuous control problems cannot be solved effectively by discretizing individual dimensions of the action space due to the exponentially large number of bins over which policies would have to be learned. In this paper, we draw inspiration from the recent success of sequence-to-sequence models for structured prediction problems to develop policies over discretized spaces. Central to this method is the realization that complex functions over high dimensional spaces can be modeled by neural networks that predict one dimension at a time. Specifically, we show how Q-values and policies over continuous spaces can be modeled using a next step prediction model over discretized dimensions. With this parameterization, it is possible to both leverage the compositional structure of action spaces during learning, as well as compute maxima over action spaces (approximately). On a simple example task we demonstrate empirically that our method can perform global search, which effectively gets around the local optimization issues that plague DDPG. We apply the technique to off-policy (Q-learning) methods and show that our method can achieve the state-of-the-art for off-policy methods on several continuous control tasks.
References in corpus (10)
- Sequence to Sequence Learning with Neural Networks
- Neural Architecture Search with Reinforcement Learning
- Weight Uncertainty in Neural Networks
- The Loss Surfaces of Multilayer Networks
- Reproducibility of Benchmarked Deep Reinforcement Learning Tasks for Continuous Control
- Equivalence Between Policy Gradients and Soft Q-Learning
- Reward Augmented Maximum Likelihood for Neural Structured Prediction
- Learning to Draw Samples: With Application to Amortized MLE for Generative Adversarial Learning
- PixelCNN++: Improving the PixelCNN with Discretized Logistic Mixture Likelihood and Other Modifications
- Towards Generalization and Simplicity in Continuous Control
Cited by in corpus (26)
- A Brief Survey of Deep Reinforcement Learning
- Hindsight Experience Replay
- Reproducibility of Benchmarked Deep Reinforcement Learning Tasks for Continuous Control
- Temporal Difference Models: Model-Free Deep RL for Model-Based Control
- A Generalized Algorithm for Multi-Objective Reinforcement Learning and Policy Adaptation
- RecSim: A Configurable Simulation Platform for Recommender Systems
- Discrete and Continuous Action Representation for Practical RL in Video Games
- Q-Learning in enormous action spaces via amortized approximate maximization
- Trust-PCL: An Off-Policy Trust Region Method for Continuous Control
- Leave no Trace: Learning to Reset for Safe and Autonomous Reinforcement Learning
- Reinforcement Learning for Slate-based Recommender Systems: A Tractable Decomposition and Practical Methodology
- Structure in Deep Reinforcement Learning: A Survey and Open Problems
- Discretizing Continuous Action Space for On-Policy Optimization
- QDP: Learning to Sequentially Optimise Quasi-Static and Dynamic Manipulation Primitives for Robotic Cloth Manipulation
- Deep Reinforcement Learning Based Volt-VAR Optimization in Smart Distribution Systems
- Tonic: A Deep Reinforcement Learning Library for Fast Prototyping and Benchmarking
- EMaQ: Expected-Max Q-Learning Operator for Simple Yet Effective Offline and Online RL
- Learning to Represent Action Values as a Hypergraph on the Action Vertices
- Continuous Control with Action Quantization from Demonstrations
- The Taxicab Sampler: MCMC for Discrete Spaces with Application to Tree Models
- Learning to Compose Hierarchical Object-Centric Controllers for Robotic Manipulation
- Simplex Decomposition for Portfolio Allocation Constraints in Reinforcement Learning
- Soft Actor-Critic With Integer Actions
- Policy learning in SE(3) action spaces
- Implicitly Regularized RL with Implicit Q-Values
- Reinforcement Learning of Implicit and Explicit Control Flow in Instructions