Automated Reinforcement Learning (AutoRL): A Survey and Open Problems
arXiv:2201.03916 · doi:10.1613/jair.1.13596
Abstract
The combination of Reinforcement Learning (RL) with deep learning has led to a series of impressive feats, with many believing (deep) RL provides a path towards generally capable agents. However, the success of RL agents is often highly sensitive to design choices in the training process, which may require tedious and error-prone manual tuning. This makes it challenging to use RL for new problems, while also limits its full potential. In many other areas of machine learning, AutoML has shown it is possible to automate such design choices and has also yielded promising initial results when applied to RL. However, Automated Reinforcement Learning (AutoRL) involves not only standard applications of AutoML but also includes additional challenges unique to RL, that naturally produce a different set of methods. As such, AutoRL has been emerging as an important area of research in RL, providing promise in a variety of applications from RNA design to playing games such as Go. Given the diversity of methods and environments considered in RL, much of the research has been conducted in distinct subfields, ranging from meta-learning to evolution. In this survey we seek to unify the field of AutoRL, we provide a common taxonomy, discuss each area in detail and pose open problems which would be of interest to researchers going forward.
Published in JAIR. Co-first authors and co-last authors are listed in alphabetical order
References in corpus (55)
- Batch Normalization: Accelerating Deep Network Training by Reducing Internal Covariate Shift
- PyTorch: An Imperative Style, High-Performance Deep Learning Library
- A Tutorial on Bayesian Optimization of Expensive Cost Functions, with Application to Active User Modeling and Hierarchical Reinforcement Learning
- Solving Rubik's Cube with a Robot Hand
- Conservative Q-Learning for Offline Reinforcement Learning
- RL: Fast Reinforcement Learning via Slow Reinforcement Learning
- Population Based Training of Neural Networks
- A Distributional Perspective on Reinforcement Learning
- Non-stochastic Best Arm Identification and Hyperparameter Optimization
- Efficient Off-Policy Meta-Reinforcement Learning via Probabilistic Context Variables
- Reproducibility of Benchmarked Deep Reinforcement Learning Tasks for Continuous Control
- In Search of the Real Inductive Bias: On the Role of Implicit Regularization in Deep Learning
- Stabilizing Transformers for Reinforcement Learning
- Learning to Utilize Shaping Rewards: A New Approach of Reward Shaping
- Revisiting Fundamentals of Experience Replay
- Bayesian Optimization in AlphaGo
- Implicit Regularization in Deep Learning
- Parallel and Distributed Thompson Sampling for Large-scale Accelerated Exploration of Chemical Space
- VariBAD: A Very Good Method for Bayes-Adaptive Deep RL via Meta-Learning
- Improving Generalization in Meta Reinforcement Learning using Learned Objectives
- Open-Ended Learning Leads to Generally Capable Agents
- Enhanced POET: Open-Ended Reinforcement Learning through Unbounded Invention of Learning Challenges and their Solutions
- Reward Shaping via Meta-Learning
- Deep Reinforcement Learning at the Edge of the Statistical Precipice
- Meta-Gradient Reinforcement Learning with an Objective Discovered Online
- Automatic Curriculum Learning through Value Disagreement
- On the Importance of Hyperparameter Optimization for Model-based Reinforcement Learning
- Emergent Complexity and Zero-shot Transfer via Unsupervised Environment Design
- Evolving Rewards to Automate Reinforcement Learning
- On Inductive Biases in Deep Reinforcement Learning
- Brax -- A Differentiable Physics Engine for Large Scale Rigid Body Simulation
- Asymmetric self-play for automatic goal discovery in robotic manipulation
- Evaluating the Performance of Reinforcement Learning Algorithms
- Observational Overfitting in Reinforcement Learning
- Think Global and Act Local: Bayesian Optimisation over High-Dimensional Categorical and Mixed Search Spaces
- D2RL: Deep Dense Architectures in Reinforcement Learning
- From Motor Control to Team Play in Simulated Humanoid Football
- MiniHack the Planet: A Sandbox for Open-Ended Reinforcement Learning Research
- An Empirical Study on Hyperparameters and their Interdependence for RL Generalization
- A Greedy Approach to Adapting the Trace Parameter for Temporal Difference Learning
- Ready Policy One: World Building Through Active Learning
- Revisiting Design Choices in Offline Model-Based Reinforcement Learning
- TempoRL: Learning When to Act
- Discovery of Options via Meta-Learned Subgoals
- Environment Generation for Zero-Shot Compositional Reinforcement Learning
- Faster Improvement Rate Population Based Training
- TeachMyAgent: a Benchmark for Automatic Curriculum Learning in Deep RL
- Quantity vs. Quality: On Hyperparameter Optimization for Deep Reinforcement Learning
- Self-Paced Context Evaluation for Contextual Reinforcement Learning
- CARL: A Benchmark for Contextual and Adaptive Reinforcement Learning
- Tuning Mixed Input Hyperparameters on the Fly for Efficient Population Based AutoRL
- Adaptive Temporal-Difference Learning for Policy Evaluation with Per-State Uncertainty Estimates
- Learning Synthetic Environments for Reinforcement Learning with Evolution Strategies
- The Cross-environment Hyperparameter Setting Benchmark for Reinforcement Learning
- Hyperparameters in Contextual RL are Highly Situational
Cited by in corpus (5)
- Evolutionary Reinforcement Learning: A Survey
- General Purpose Artificial Intelligence Systems (GPAIS): Properties, Definition, Taxonomy, Societal Implications and Responsible Governance
- Sim2real for Autonomous Vehicle Control using Executable Digital Twin
- Structure in Deep Reinforcement Learning: A Survey and Open Problems
- On the Importance of Reward Design in Reinforcement Learning-based Dynamic Algorithm Configuration: A Case Study on OneMax with (1+(,))-GA