Dopamine: A Research Framework for Deep Reinforcement Learning
arXiv:1812.06110
Abstract
Deep reinforcement learning (deep RL) research has grown significantly in recent years. A number of software offerings now exist that provide stable, comprehensive implementations for benchmarking. At the same time, recent deep RL research has become more diverse in its goals. In this paper we introduce Dopamine, a new research framework for deep RL that aims to support some of that diversity. Dopamine is open-source, TensorFlow-based, and provides compact and reliable implementations of some state-of-the-art deep RL agents. We complement this offering with a taxonomy of the different research objectives in deep RL research. While by no means exhaustive, our analysis highlights the heterogeneity of research in the field, and the value of frameworks such as ours.
Cited by in corpus (20)
- Soft Actor-Critic for Discrete Action Settings
- Benchmarking Batch Deep Reinforcement Learning Algorithms
- DeepMDP: Learning Continuous Latent Space Models for Representation Learning
- Hyperbolic Discounting and Learning over Multiple Horizons
- rlpyt: A Research Code Base for Deep Reinforcement Learning in PyTorch
- RecSim: A Configurable Simulation Platform for Recommender Systems
- AllenAct: A Framework for Embodied AI Research
- Measuring the Reliability of Reinforcement Learning Algorithms
- SEED RL: Scalable and Efficient Deep-RL with Accelerated Central Inference
- A Geometric Perspective on Optimal Representations for Reinforcement Learning
- Reinforcement Learning for Slate-based Recommender Systems: A Tractable Decomposition and Practical Methodology
- DOM-Q-NET: Grounded RL on Structured Language
- MULEX: Disentangling Exploitation from Exploration in Deep RL
- VALAN: Vision and Language Agent Navigation
- Catalyst.RL: A Distributed Framework for Reproducible RL Research
- Lessons from Contextual Bandit Learning in a Customer Support Bot
- Sample Efficient Ensemble Learning with Catalyst.RL
- TDprop: Does Jacobi Preconditioning Help Temporal Difference Learning?
- Fair treatment allocations in social networks
- ConQUR: Mitigating Delusional Bias in Deep Q-learning