Continuous Control with Action Quantization from Demonstrations
arXiv:2110.10149
Abstract
In this paper, we propose a novel Reinforcement Learning (RL) framework for problems with continuous action spaces: Action Quantization from Demonstrations (AQuaDem). The proposed approach consists in learning a discretization of continuous action spaces from human demonstrations. This discretization returns a set of plausible actions (in light of the demonstrations) for each input state, thus capturing the priors of the demonstrator and their multimodal behavior. By discretizing the action space, any discrete action deep RL technique can be readily applied to the continuous control problem. Experiments show that the proposed approach outperforms state-of-the-art methods such as SAC in the RL setup, and GAIL in the Imitation Learning setup. We provide a website with interactive videos: https://google-research.github.io/aquadem/ and make the code available: https://github.com/google-research/google-research/tree/master/aquadem.
Accepted to ICML 2022
References in corpus (21)
- Soft Actor-Critic: Off-Policy Maximum Entropy Deep Reinforcement Learning with a Stochastic Actor
- Dota 2 with Large Scale Deep Reinforcement Learning
- QT-Opt: Scalable Deep Reinforcement Learning for Vision-Based Robotic Manipulation
- Leveraging Demonstrations for Deep Reinforcement Learning on Robotics Problems with Sparse Rewards
- Deep Q-learning from Demonstrations
- Behavior Regularized Offline Reinforcement Learning
- A Minimalist Approach to Offline Reinforcement Learning
- Learning Montezuma's Revenge from a Single Demonstration
- What Matters in Learning from Offline Human Demonstrations for Robot Manipulation
- Acme: A Research Framework for Distributed Reinforcement Learning
- Discriminator-Actor-Critic: Addressing Sample Inefficiency and Reward Bias in Adversarial Imitation Learning
- Discrete Sequential Prediction of Continuous Actions for Deep RL
- Munchausen Reinforcement Learning
- What Matters for Adversarial Imitation Learning?
- Distributional Policy Optimization: An Alternative Approach for Continuous Control
- Random Expert Distillation: Imitation Learning via Expert Policy Support Estimation
- Q-Learning for Continuous Actions with Cross-Entropy Guided Policies
- Hyperparameter Selection for Imitation Learning
- Deep Radial-Basis Value Functions for Continuous Control
- Implicit Behavioral Cloning
- Implicitly Regularized RL with Implicit Q-Values