TensorFlow Agents: Efficient Batched Reinforcement Learning in TensorFlow
arXiv:1709.02878
Abstract
We introduce TensorFlow Agents, an efficient infrastructure paradigm for building parallel reinforcement learning algorithms in TensorFlow. We simulate multiple environments in parallel, and group them to perform the neural network computation on a batch rather than individual observations. This allows the TensorFlow execution engine to parallelize computation, without the need for manual synchronization. Environments are stepped in separate Python processes to progress them in parallel without interference of the global interpreter lock. As part of this project, we introduce BatchPPO, an efficient implementation of the proximal policy optimization algorithm. By open sourcing TensorFlow Agents, we hope to provide a flexible starting point for future projects that accelerates future research in the field.
White paper, 7 pages
References in corpus (4)
Cited by in corpus (8)
- Dopamine: A Research Framework for Deep Reinforcement Learning
- Lyapunov-based Safe Policy Optimization for Continuous Control
- Sim-to-Real: Learning Agile Locomotion For Quadruped Robots
- Graph Policy Gradients for Large Scale Robot Control
- RLgraph: Modular Computation Graphs for Deep Reinforcement Learning
- SURREAL-System: Fully-Integrated Stack for Distributed Deep Reinforcement Learning
- Autonomous Reinforcement Learning via Subgoal Curricula
- WALL-E: An Efficient Reinforcement Learning Research Framework