Challenges of Real-World Reinforcement Learning
arXiv:1904.12901
Abstract
Reinforcement learning (RL) has proven its worth in a series of artificial domains, and is beginning to show some successes in real-world scenarios. However, much of the research advances in RL are often hard to leverage in real-world systems due to a series of assumptions that are rarely satisfied in practice. We present a set of nine unique challenges that must be addressed to productionize RL to real world problems. For each of these challenges, we specify the exact meaning of the challenge, present some approaches from the literature, and specify some metrics for evaluating that challenge. An approach that addresses all nine challenges would be applicable to a large number of real world problems. We also present an example domain that has been modified to present these challenges as a testbed for practical RL research.
References in corpus (10)
- Methods for Interpreting and Understanding Deep Neural Networks
- DeepMind Control Suite
- Doubly Robust Policy Evaluation and Learning
- Safe Exploration in Continuous Action Spaces
- Hybrid Reward Architecture for Reinforcement Learning
- Policy Gradients with Variance Related Risk Criteria
- Policy Gradient for Coherent Risk Measures
- Value constrained model-free continuous control
- Deep Robust Kalman Filter
- Situational Awareness by Risk-Conscious Skills
Cited by in corpus (37)
- Language-guided Navigation via Cross-Modal Grounding and Alternate Adversarial Learning
- RecSim: A Configurable Simulation Platform for Recommender Systems
- ABIDES-Gym: Gym Environments for Multi-Agent Discrete Event Simulation and Application to Financial Markets
- Representation Matters: Offline Pretraining for Sequential Decision Making
- IPO: Interior-point Policy Optimization under Constraints
- Risk-Averse Offline Reinforcement Learning
- D2RL: Deep Dense Architectures in Reinforcement Learning
- Challenges of Applying Deep Reinforcement Learning in Dynamic Dispatching
- Diverse Auto-Curriculum is Critical for Successful Real-World Multiagent Learning Systems
- GST: Group-Sparse Training for Accelerating Deep Reinforcement Learning
- RecSim NG: Toward Principled Uncertainty Modeling for Recommender Ecosystems
- The Benchmark Lottery
- Self-Imitation Advantage Learning
- A Framework for Studying Reinforcement Learning and Sim-to-Real in Robot Soccer
- Balancing Constraints and Rewards with Meta-Gradient D4PG
- Multi-Objective SPIBB: Seldonian Offline Policy Improvement with Safety Constraints in Finite MDPs
- Linear Representation Meta-Reinforcement Learning for Instant Adaptation
- Evaluating the progress of Deep Reinforcement Learning in the real world: aligning domain-agnostic and domain-specific research
- Combining Pessimism with Optimism for Robust and Efficient Model-Based Deep Reinforcement Learning
- Energy and Thermal-aware Resource Management of Cloud Data Centres: A Taxonomy and Future Directions
- Explaining Conditions for Reinforcement Learning Behaviors from Real and Imagined Data
- Quick Learner Automated Vehicle Adapting its Roadmanship to Varying Traffic Cultures with Meta Reinforcement Learning
- World Programs for Model-Based Learning and Planning in Compositional State and Action Spaces
- Boosting Offline Reinforcement Learning with Residual Generative Modeling
- The Big Three: A Methodology to Increase Data Science ROI by Answering the Questions Companies Care About
- Measuring Progress in Deep Reinforcement Learning Sample Efficiency
- Reducing Conservativeness Oriented Offline Reinforcement Learning
- A survey of benchmarking frameworks for reinforcement learning
- Bellman: A Toolbox for Model-Based Reinforcement Learning in TensorFlow
- Provable Multi-Objective Reinforcement Learning with Generative Models
- Can Reinforcement Learning for Continuous Control Generalize Across Physics Engines?
- How Do You Act? An Empirical Study to Understand Behavior of Deep Reinforcement Learning Agents
- Learning robust driving policies without online exploration
- Learning Robust Controllers Via Probabilistic Model-Based Policy Search
- Robust Reinforcement Learning under model misspecification
- Resonance: Replacing Software Constants with Context-Aware Models in Real-time Communication
- GrowSpace: Learning How to Shape Plants