An empirical investigation of the challenges of real-world reinforcement learning
arXiv:2003.11881
Abstract
Reinforcement learning (RL) has proven its worth in a series of artificial domains, and is beginning to show some successes in real-world scenarios. However, much of the research advances in RL are hard to leverage in real-world systems due to a series of assumptions that are rarely satisfied in practice. In this work, we identify and formalize a series of independent challenges that embody the difficulties that must be addressed for RL to be commonly deployed in real-world systems. For each challenge, we define it formally in the context of a Markov Decision Process, analyze the effects of the challenge on state-of-the-art learning algorithms, and present some existing attempts at tackling it. We believe that an approach that addresses our set of proposed challenges would be readily deployable in a large number of real world problems. Our proposed challenges are implemented in a suite of continuous control environments called the realworldrl-suite which we propose an as an open-source benchmark.
arXiv admin note: text overlap with arXiv:1904.12901
References in corpus (29)
- Offline Reinforcement Learning: Tutorial, Review, and Perspectives on Open Problems
- Emergence of Locomotion Behaviours in Rich Environments
- DeepMind Control Suite
- Safe Exploration in Continuous Action Spaces
- Challenges of Real-World Reinforcement Learning
- Behavior Regularized Offline Reinforcement Learning
- MOPO: Model-based Offline Policy Optimization
- Hybrid Reward Architecture for Reinforcement Learning
- Advantage-Weighted Regression: Simple and Scalable Off-Policy Reinforcement Learning
- MOReL : Model-Based Offline Reinforcement Learning
- Way Off-Policy Batch Deep Reinforcement Learning of Implicit Human Preferences in Dialog
- Constrained Policy Optimization
- Critic Regularized Regression
- Deep Dynamics Models for Learning Dexterous Manipulation
- Acme: A Research Framework for Distributed Reinforcement Learning
- RecSim: A Configurable Simulation Platform for Recommender Systems
- Deep Online Learning via Meta-Learning: Continual Adaptation for Model-Based RL
- Responsive Safety in Reinforcement Learning by PID Lagrangian Methods
- Policy Gradient for Coherent Risk Measures
- Value constrained model-free continuous control
- A Distributional View on Multi-Objective Policy Optimization
- A Bayesian Approach to Robust Reinforcement Learning
- Model-Based Offline Planning
- Deep Robust Kalman Filter
- Real-Time Reinforcement Learning
- On Ensuring that Intelligent Machines Are Well-Behaved
- Balancing Constraints and Rewards with Meta-Gradient D4PG
- Robust Constrained Reinforcement Learning for Continuous Control with Model Misspecification
- Situational Awareness by Risk-Conscious Skills
Cited by in corpus (17)
- D4RL: Datasets for Deep Data-Driven Reinforcement Learning
- A Survey of Zero-shot Generalisation in Deep Reinforcement Learning
- Pretraining Representations for Data-Efficient Reinforcement Learning
- A Self-Tuning Actor-Critic Algorithm
- CausalWorld: A Robotic Manipulation Benchmark for Causal Structure and Transfer Learning
- NeoRL: A Near Real-World Benchmark for Offline Reinforcement Learning
- Habitat 2.0: Training Home Assistants to Rearrange their Habitat
- Model-Based Offline Planning
- On the Emergence of Whole-body Strategies from Humanoid Robot Push-recovery Learning
- D2RL: Deep Dense Architectures in Reinforcement Learning
- Learning and Planning in Complex Action Spaces
- Diverse Auto-Curriculum is Critical for Successful Real-World Multiagent Learning Systems
- Balancing Constraints and Rewards with Meta-Gradient D4PG
- Robust Constrained Reinforcement Learning for Continuous Control with Model Misspecification
- Is Bang-Bang Control All You Need? Solving Continuous Control with Bernoulli Policies
- Monte Carlo Tree Search for high precision manufacturing
- GrowSpace: Learning How to Shape Plants