Data-efficient Deep Reinforcement Learning for Dexterous Manipulation
arXiv:1704.03073
Abstract
Deep learning and reinforcement learning methods have recently been used to solve a variety of problems in continuous control domains. An obvious application of these techniques is dexterous manipulation tasks in robotics which are difficult to solve using traditional control theory or hand-engineered approaches. One example of such a task is to grasp an object and precisely stack it on another. Solving this difficult and practically relevant problem in the real world is an important long-term goal for the field of robotics. Here we take a step towards this goal by examining the problem in simulation and providing models and techniques aimed at solving it. We introduce two extensions to the Deep Deterministic Policy Gradient algorithm (DDPG), a model-free Q-learning based method, which make it significantly more data-efficient and scalable. Our results show that by making extensive use of off-policy data and replay, it is possible to find control policies that robustly grasp objects and stack them. Further, our results hint that it may soon be feasible to train successful stacking policies by collecting interactions on real robots.
12 pages, 5 Figures
References in corpus (1)
Cited by in corpus (43)
- Soft Actor-Critic: Off-Policy Maximum Entropy Deep Reinforcement Learning with a Stochastic Actor
- Addressing Function Approximation Error in Actor-Critic Methods
- StarCraft II: A New Challenge for Reinforcement Learning
- Hindsight Experience Replay
- Distributed Distributional Deterministic Policy Gradients
- Transferring End-to-End Visuomotor Control from Simulation to Real World for a Multi-Stage Task
- Reinforcement and Imitation Learning for Diverse Visuomotor Skills
- Reinforcement Learning for Robotic Manipulation using Simulated Locomotion Demonstrations
- Value constrained model-free continuous control
- Sim-to-Real Transfer of Accurate Grasping with Eye-In-Hand Observations and Continuous Control
- TriFinger: An Open-Source Robot for Learning Dexterity
- Bayesian policy selection using active inference
- Asymmetric self-play for automatic goal discovery in robotic manipulation
- Towards Practical Multi-Object Manipulation using Relational Reinforcement Learning
- Whole-Body Control of a Mobile Manipulator using End-to-End Reinforcement Learning
- What Matters for Adversarial Imitation Learning?
- Deep Reinforcement Learning for Dexterous Manipulation with Concept Networks
- Goal-Auxiliary Actor-Critic for 6D Robotic Grasping with Point Clouds
- Form2Fit: Learning Shape Priors for Generalizable Assembly from Disassembly
- Curiosity-Driven Multi-Criteria Hindsight Experience Replay
- ConfuciuX: Autonomous Hardware Resource Assignment for DNN Accelerators using Reinforcement Learning
- C-Learning: Learning to Achieve Goals via Recursive Classification
- Acceleration of Actor-Critic Deep Reinforcement Learning for Visual Grasping in Clutter by State Representation Learning Based on Disentanglement of a Raw Input Image
- Local Search for Policy Iteration in Continuous Control
- How to Close Sim-Real Gap? Transfer with Segmentation!
- Physically Embedded Planning Problems: New Challenges for Reinforcement Learning
- Disentangled Planning and Control in Vision Based Robotics via Reward Machines
- Benchmarking Structured Policies and Policy Optimization for Real-World Dexterous Object Manipulation
- Exploring Restart Distributions
- Competitive Experience Replay
- Beyond Tabula-Rasa: a Modular Reinforcement Learning Approach for Physically Embedded 3D Sokoban
- Imaginary Hindsight Experience Replay: Curious Model-based Learning for Sparse Reward Tasks
- Adversarial Learning of Task-Oriented Neural Dialog Models
- Follow the Object: Curriculum Learning for Manipulation Tasks with Imagined Goals
- Efficient Robotic Object Search via HIEM: Hierarchical Policy Learning with Intrinsic-Extrinsic Modeling
- Expanding Motor Skills through Relay Neural Networks
- Regularly Updated Deterministic Policy Gradient Algorithm
- How Do You Act? An Empirical Study to Understand Behavior of Deep Reinforcement Learning Agents
- Deep Reinforcement Learning for Tactile Robotics: Learning to Type on a Braille Keyboard
- Generative Exploration and Exploitation
- Contrastive Active Inference
- Parallelized Reverse Curriculum Generation
- ACDER: Augmented Curiosity-Driven Experience Replay