Solving Rubik's Cube with a Robot Hand
arXiv:1910.07113
Abstract
We demonstrate that models trained only in simulation can be used to solve a manipulation problem of unprecedented complexity on a real robot. This is made possible by two key components: a novel algorithm, which we call automatic domain randomization (ADR) and a robot platform built for machine learning. ADR automatically generates a distribution over randomized environments of ever-increasing difficulty. Control policies and vision state estimators trained with ADR exhibit vastly improved sim2real transfer. For control policies, memory-augmented models trained on an ADR-generated distribution of environments show clear signs of emergent meta-learning at test time. The combination of ADR with our custom robot platform allows us to solve a Rubik's cube with a humanoid robot hand, which involves both control and state estimation problems. Videos summarizing our results are available: https://openai.com/blog/solving-rubiks-cube/
References in corpus (14)
- Learning agile and dynamic motor skills for legged robots
- Mastering Chess and Shogi by Self-Play with a General Reinforcement Learning Algorithm
- Robots that can adapt like animals
- Robust Adversarial Reinforcement Learning
- Learning to reinforcement learn
- Domain Randomization for Transferring Deep Neural Networks from Simulation to the Real World
- Learning Invariant Feature Spaces to Transfer Skills with Reinforcement Learning
- Deep Dynamics Models for Learning Dexterous Manipulation
- Distilling Policy Distillation
- Reinforcement Learning for Pivoting Task
- Active Domain Randomization
- Concurrent Meta Reinforcement Learning
- ORRB -- OpenAI Remote Rendering Backend
- First Experiments with PowerPlay
Cited by in corpus (17)
- Multi-Task Learning with Deep Neural Networks: A Survey
- The Unreasonable Effectiveness of Deep Learning in Artificial Intelligence
- The Next Decade in AI: Four Steps Towards Robust Artificial Intelligence
- Phasic Policy Gradient
- Generative Language Modeling for Automated Theorem Proving
- Whole-Body Control of a Mobile Manipulator using End-to-End Reinforcement Learning
- Revisiting Design Choices in Proximal Policy Optimization
- Traversing the Reality Gap via Simulator Tuning
- Curriculum in Gradient-Based Meta-Reinforcement Learning
- Ensuring Monotonic Policy Improvement in Entropy-regularized Value-based Reinforcement Learning
- Scalable sim-to-real transfer of soft robot designs
- Self-Adapting Recurrent Models for Object Pushing from Learning in Simulation
- Deep Reinforcement Learning with Linear Quadratic Regulator Regions
- Scene Graph Modification Based on Natural Language Commands
- Predicting Sim-to-Real Transfer with Probabilistic Dynamics Models
- Deep Reinforcement Learning for Tactile Robotics: Learning to Type on a Braille Keyboard
- RL agents Implicitly Learning Human Preferences