Reset-free Trial-and-Error Learning for Robot Damage Recovery
arXiv:1610.04213 · doi:10.1016/j.robot.2017.11.010
Abstract
The high probability of hardware failures prevents many advanced robots (e.g., legged robots) from being confidently deployed in real-world situations (e.g., post-disaster rescue). Instead of attempting to diagnose the failures, robots could adapt by trial-and-error in order to be able to complete their tasks. In this situation, damage recovery can be seen as a Reinforcement Learning (RL) problem. However, the best RL algorithms for robotics require the robot and the environment to be reset to an initial state after each episode, that is, the robot is not learning autonomously. In addition, most of the RL methods for robotics do not scale well with complex robots (e.g., walking robots) and either cannot be used at all or take too long to converge to a solution (e.g., hours of learning). In this paper, we introduce a novel learning algorithm called "Reset-free Trial-and-Error" (RTE) that (1) breaks the complexity by pre-generating hundreds of possible behaviors with a dynamics simulator of the intact robot, and (2) allows complex robots to quickly recover from damage while completing their tasks and taking the environment into account. We evaluate our algorithm on a simulated wheeled robot, a simulated six-legged robot, and a real six-legged walking robot that are damaged in several ways (e.g., a missing leg, a shortened leg, faulty motor, etc.) and whose objective is to reach a sequence of targets in an arena. Our experiments show that the robots can recover most of their locomotion abilities in an environment with obstacles, and without any human intervention.
18 pages, 16 figures, 3 tables, 6 pseudocodes/algorithms, video at https://youtu.be/IqtyHFrb3BU, code at https://github.com/resibots/chatzilygeroudis_2018_rte
References in corpus (10)
- Trust Region Policy Optimization
- Dropout as a Bayesian Approximation: Representing Model Uncertainty in Deep Learning
- Robots that can adapt like animals
- Gaussian Processes for Data-Efficient Learning in Robotics and Control
- Active Learning of Inverse Models with Intrinsically Motivated Goal Exploration in Robots
- Illuminating search spaces by mapping elites
- Multiple chaotic central pattern generators with learning for legged locomotion and malfunction compensation
- Patchwork Kriging for Large-scale Gaussian Process Regression
- Limbo: A Fast and Flexible Library for Bayesian Optimization
- Reset-Free Guided Policy Search: Efficient Deep Reinforcement Learning with Stochastic Initial States
Cited by in corpus (28)
- A soft robot that adapts to environments through shape change
- Learning to Walk in the Real World with Minimal Human Effort
- Discovering the Elite Hypervolume by Leveraging Interspecies Correlation
- Quality Diversity for Multi-task Optimization
- Scaling MAP-Elites to Deep Neuroevolution
- Automated shapeshifting for function recovery in damaged robots
- Never Stop Learning: The Effectiveness of Fine-Tuning in Robotic Reinforcement Learning
- Hierarchical Behavioral Repertoires with Unsupervised Descriptors
- Adaptive Prior Selection for Repertoire-based Online Adaptation in Robotics
- From exploration to control: learning object manipulation skills through novelty search and local adaptation
- Dynamics-Aware Quality-Diversity for Efficient Learning of Skill Repertoires
- Evolving the Behavior of Machines: From Micro to Macroevolution
- Hierarchical Quality-Diversity for Online Damage Recovery
- Ecological Reinforcement Learning
- Map-based Multi-Policy Reinforcement Learning: Enhancing Adaptability of Robots by Deep Reinforcement Learning
- Learning to Walk Autonomously via Reset-Free Quality-Diversity
- Using Centroidal Voronoi Tessellations to Scale Up the Multi-dimensional Archive of Phenotypic Elites Algorithm
- Regenerating Soft Robots through Neural Cellular Automata
- Don't Bet on Luck Alone: Enhancing Behavioral Reproducibility of Quality-Diversity Solutions in Uncertain Domains
- Safety-Aware Robot Damage Recovery Using Constrained Bayesian Optimization and Simulated Priors
- Relevance-guided Unsupervised Discovery of Abilities with Quality-Diversity Algorithms
- Fault-Aware Robust Control via Adversarial Reinforcement Learning
- Black-Box Data-efficient Policy Search for Robotics
- Enhancing MAP-Elites with Multiple Parallel Evolution Strategies
- A Simple Approach to Continual Learning by Transferring Skill Parameters
- Graph-Based Controller Synthesis for Safety-Constrained, Resilient Systems
- Learning in Sparse Rewards settings through Quality-Diversity algorithms
- QED: using Quality-Environment-Diversity to evolve resilient robot swarms