Scalable trust-region method for deep reinforcement learning using Kronecker-factored approximation
arXiv:1708.05144
Abstract
In this work, we propose to apply trust region optimization to deep reinforcement learning using a recently proposed Kronecker-factored approximation to the curvature. We extend the framework of natural policy gradient and propose to optimize both the actor and the critic using Kronecker-factored approximate curvature (K-FAC) with trust region; hence we call our method Actor Critic using Kronecker-Factored Trust Region (ACKTR). To the best of our knowledge, this is the first scalable trust region natural gradient method for actor-critic methods. It is also a method that learns non-trivial tasks in continuous control as well as discrete control policies directly from raw pixel inputs. We tested our approach across discrete domains in Atari games as well as continuous domains in the MuJoCo environment. With the proposed methods, we are able to achieve higher rewards and a 2- to 3-fold improvement in sample efficiency on average, compared to previous state-of-the-art on-policy actor-critic methods. Code is available at https://github.com/openai/baselines
14 pages, 9 figures; update github repo link
References in corpus (1)
Cited by in corpus (35)
- Soft Actor-Critic for Discrete Action Settings
- A Survey of Deep Reinforcement Learning in Video Games
- Phasic Policy Gradient
- Dealing with Sparse Rewards in Reinforcement Learning
- AllenAct: A Framework for Embodied AI Research
- Reinforcement Learning with Competitive Ensembles of Information-Constrained Primitives
- A Machine Learning Approach to Routing
- Let's Play Again: Variability of Deep Reinforcement Learning Agents in Atari Environments
- ConfuciuX: Autonomous Hardware Resource Assignment for DNN Accelerators using Reinforcement Learning
- Finding and Visualizing Weaknesses of Deep Reinforcement Learning Agents
- Towards Similarity Graphs Constructed by Deep Reinforcement Learning
- Kalman meets Bellman: Improving Policy Evaluation through Value Tracking
- ROS2Learn: a reinforcement learning framework for ROS 2
- Scalable Multi-Agent Inverse Reinforcement Learning via Actor-Attention-Critic
- Toybox: A Suite of Environments for Experimental Evaluation of Deep Reinforcement Learning
- Partially Detected Intelligent Traffic Signal Control: Environmental Adaptation
- Trust Region Value Optimization using Kalman Filtering
- Off-Policy Actor-Critic in an Ensemble: Achieving Maximum General Entropy and Effective Environment Exploration in Deep Reinforcement Learning
- On-Policy Trust Region Policy Optimisation with Replay Buffers
- Reinforcement Learning with Structured Hierarchical Grammar Representations of Actions
- RL STaR Platform: Reinforcement Learning for Simulation based Training of Robots
- Disentangling Dynamics and Returns: Value Function Decomposition with Future Prediction
- LeRoP: A Learning-Based Modular Robot Photography Framework
- An Empirical Analysis of Proximal Policy Optimization with Kronecker-factored Natural Gradients
- Recurrent Value Functions
- D2C 2.0: Decoupled Data-Based Approach for Learning to Control Stochastic Nonlinear Systems via Model-Free ILQR
- Scalable and Practical Natural Gradient for Large-Scale Deep Learning
- Monte-Carlo Tree Search for Policy Optimization
- Biologically inspired architectures for sample-efficient deep reinforcement learning
- Which Channel to Ask My Question? Personalized Customer Service Request Stream Routing using Deep Reinforcement Learning
- Quasi-Newton Trust Region Policy Optimization
- Genetic-Gated Networks for Deep Reinforcement
- Rogue-Gym: A New Challenge for Generalization in Reinforcement Learning
- Regularly Updated Deterministic Policy Gradient Algorithm
- Deep Active Localization