Continuous-Discrete Reinforcement Learning for Hybrid Control in Robotics
arXiv:2001.00449
Abstract
Many real-world control problems involve both discrete decision variables - such as the choice of control modes, gear switching or digital outputs - as well as continuous decision variables - such as velocity setpoints, control gains or analogue outputs. However, when defining the corresponding optimal control or reinforcement learning problem, it is commonly approximated with fully continuous or fully discrete action spaces. These simplifications aim at tailoring the problem to a particular algorithm or solver which may only support one type of action space. Alternatively, expert heuristics are used to remove discrete actions from an otherwise continuous space. In contrast, we propose to treat hybrid problems in their 'native' form by solving them with hybrid reinforcement learning, which optimizes for discrete and continuous actions simultaneously. In our experiments, we first demonstrate that the proposed approach efficiently solves such natively hybrid reinforcement learning problems. We then show, both in simulation and on robotic hardware, the benefits of removing possibly imperfect expert-designed heuristics. Lastly, hybrid reinforcement learning encourages us to rethink problem definitions. We propose reformulating control problems, e.g. by adding meta actions, to improve exploration or reduce mechanical wear and tear.
Presented at the 3rd Conference on Robot Learning (CoRL 2019), Osaka, Japan. Video: https://youtu.be/eUqQDLQXb7I
References in corpus (12)
- Adam: A Method for Stochastic Optimization
- Continuous control with deep reinforcement learning
- Soft Actor-Critic Algorithms and Applications
- Evolution Strategies as a Scalable Alternative to Reinforcement Learning
- DeepMind Control Suite
- Parametrized Deep Q-Networks Learning: Reinforcement Learning with Discrete-Continuous Hybrid Action Space
- Reinforcement Learning with Parameterized Actions
- Deep Reinforcement Learning in Parameterized Action Space
- Multi-Pass Q-Networks for Deep Reinforcement Learning with Parameterised Action Spaces
- Relative Entropy Regularized Policy Iteration
- The Termination Critic
- Hybrid Actor-Critic Reinforcement Learning in Parameterized Action Space
Cited by in corpus (7)
- One model Packs Thousands of Items with Recurrent Conditional Query Learning
- Learning Event-triggered Control from Data through Joint Optimization
- Smooth Exploration for Robotic Reinforcement Learning
- Noise-Robust End-to-End Quantum Control using Deep Autoregressive Policy Networks
- Safety Filtering for Reinforcement Learning-based Adaptive Cruise Control
- "What, not how": Solving an under-actuated insertion task from scratch
- Decoupled Exploration and Exploitation Policies for Sample-Efficient Reinforcement Learning