Sample Efficient Actor-Critic with Experience Replay
arXiv:1611.01224
Abstract
This paper presents an actor-critic deep reinforcement learning agent with experience replay that is stable, sample efficient, and performs remarkably well on challenging environments, including the discrete 57-game Atari domain and several continuous control problems. To achieve this, the paper introduces several innovations, including truncated importance sampling with bias correction, stochastic dueling network architectures, and a new trust region policy optimization method.
20 pages. Prepared for ICLR 2017
References in corpus (3)
Cited by in corpus (98)
- Applications of Deep Learning and Reinforcement Learning to Biological Data
- Offline Reinforcement Learning: Tutorial, Review, and Perspectives on Open Problems
- A Survey and Critique of Multiagent Deep Reinforcement Learning
- Learning to Evade Static PE Machine Learning Malware Models via Reinforcement Learning
- A Deeper Look at Experience Replay
- Simple random search provides a competitive approach to reinforcement learning
- A Survey of Deep Reinforcement Learning in Video Games
- Reinforcement Learning Algorithms: An Overview and Classification
- Controlling an Autonomous Vehicle with Deep Reinforcement Learning
- GEP-PG: Decoupling Exploration and Exploitation in Deep Reinforcement Learning Algorithms
- Actor-Critic Method for High Dimensional Static Hamilton--Jacobi--Bellman Partial Differential Equations based on Neural Networks
- Automatic Data Augmentation for Generalization in Deep Reinforcement Learning
- A Survey on Deep Learning Methods for Robot Vision
- Phasic Policy Gradient
- A reinforcement learning approach to rare trajectory sampling
- Learning Implicit Credit Assignment for Cooperative Multi-Agent Reinforcement Learning
- Off-Policy Policy Gradient with State Distribution Correction
- The Mirage of Action-Dependent Baselines in Reinforcement Learning
- Boosting Soft Actor-Critic: Emphasizing Recent Experience without Forgetting the Past
- A Finite Time Analysis of Two Time-Scale Actor Critic Methods
- A Benchmarking Environment for Reinforcement Learning Based Task Oriented Dialogue Management
- Worst Cases Policy Gradients
- Prioritized Sequence Experience Replay
- Pretraining Deep Actor-Critic Reinforcement Learning Algorithms With Expert Demonstrations
- Trust-PCL: An Off-Policy Trust Region Method for Continuous Control
- An Adaptive Clipping Approach for Proximal Policy Optimization
- Recall Traces: Backtracking Models for Efficient Reinforcement Learning
- Sample-Efficient Imitation Learning via Generative Adversarial Nets
- Experience Replay with Likelihood-free Importance Weights
- Enhancing Text-based Reinforcement Learning Agents with Commonsense Knowledge
- Lipschitzness Is All You Need To Tame Off-policy Generative Adversarial Imitation Learning
- Towards Cooperation in Sequential Prisoner's Dilemmas: a Deep Multiagent Reinforcement Learning Approach
- Value Propagation Networks
- Experience Replay Using Transition Sequences
- To Learn or Not to Learn: Analyzing the Role of Learning for Navigation in Virtual Environments
- A State Aggregation Approach for Solving Knapsack Problem with Deep Reinforcement Learning
- Online Off-policy Prediction
- Deep Reinforcement Learning for Autonomous Internet of Things: Model, Applications and Challenges
- Striving for Simplicity and Performance in Off-Policy DRL: Output Normalization and Non-Uniform Sampling
- Muesli: Combining Improvements in Policy Optimization
- Double Prioritized State Recycled Experience Replay
- Deep reinforcement learning driven inspection and maintenance planning under incomplete information and constraints
- Optimizing Sequential Medical Treatments with Auto-Encoding Heuristic Search in POMDPs
- Diversity Actor-Critic: Sample-Aware Entropy Regularization for Sample-Efficient Exploration
- Direct and indirect reinforcement learning
- Effective Exploration for Deep Reinforcement Learning via Bootstrapped Q-Ensembles under Tsallis Entropy Regularization
- ROS2Learn: a reinforcement learning framework for ROS 2
- From Few to More: Large-scale Dynamic Multiagent Curriculum Learning
- Dependency-aware Attention Control for Unconstrained Face Recognition with Image Sets
- Deep Reinforcement Learning for Demand Driven Services in Logistics and Transportation Systems: A Survey
- Provably Convergent Two-Timescale Off-Policy Actor-Critic with Function Approximation
- Action Guidance: Getting the Best of Sparse Rewards and Shaped Rewards for Real-time Strategy Games
- Implicit Distributional Reinforcement Learning
- Generalized Off-Policy Actor-Critic
- Is the Policy Gradient a Gradient?
- On-Policy Trust Region Policy Optimisation with Replay Buffers
- P3O: Policy-on Policy-off Policy Optimization
- Learning What to Memorize: Using Intrinsic Motivation to Form Useful Memory in Partially Observable Reinforcement Learning
- AMBER: Adaptive Multi-Batch Experience Replay for Continuous Action Control
- An Off-policy Policy Gradient Theorem Using Emphatic Weightings
- Stealing Deep Reinforcement Learning Models for Fun and Profit
- Rethinking Supervised Learning and Reinforcement Learning in Task-Oriented Dialogue Systems
- Auto-Agent-Distiller: Towards Efficient Deep Reinforcement Learning Agents via Neural Architecture Search
- Integrating LEO Satellites and Multi-UAV Reinforcement Learning for Hybrid FSO/RF Non-Terrestrial Networks
- Review, Analysis and Design of a Comprehensive Deep Reinforcement Learning Framework
- Transferable Cost-Aware Security Policy Implementation for Malware Detection Using Deep Reinforcement Learning
- Policy Search by Target Distribution Learning for Continuous Control
- The act of remembering: a study in partially observable reinforcement learning
- The Actor-Advisor: Policy Gradient With Off-Policy Advice
- Competitive Experience Replay
- Improved Soft Actor-Critic: Mixing Prioritized Off-Policy Samples with On-Policy Experience
- Policy Optimization with Model-based Explorations
- A Dual Memory Structure for Efficient Use of Replay Memory in Deep Reinforcement Learning
- Zeroth-Order Supervised Policy Improvement
- Direct Advantage Estimation
- Fairness in Multi-agent Reinforcement Learning for Stock Trading
- Amplifying the Imitation Effect for Reinforcement Learning of UCAV's Mission Execution
- Diluted Near-Optimal Expert Demonstrations for Guiding Dialogue Stochastic Policy Optimisation
- IMPACT: Importance Weighted Asynchronous Architectures with Clipped Target Networks
- Discrete Action On-Policy Learning with Action-Value Critic
- Semi-On-Policy Training for Sample Efficient Multi-Agent Policy Gradients
- An Active Learning Framework for Efficient Robust Policy Search
- Parameter-Based Value Functions
- GAPLE: Generalizable Approaching Policy LEarning for Robotic Object Searching in Indoor Environment
- Improving the sample-efficiency of neural architecture search with reinforcement learning
- Deep-Reinforcement-Learning for Gliding and Perching Bodies
- Dr Jekyll and Mr Hyde: the Strange Case of Off-Policy Policy Updates
- Discrete-to-Deep Supervised Policy Learning
- Measuring Progress in Deep Reinforcement Learning Sample Efficiency
- An advantage actor-critic algorithm for robotic motion planning in dense and dynamic scenarios
- Lucid Dreaming for Experience Replay: Refreshing Past States with the Current Policy
- Implications of Human Irrationality for Reinforcement Learning
- On The Transferability of Deep-Q Networks
- ACDER: Augmented Curiosity-Driven Experience Replay
- Ranking Policy Gradient
- Dynamic Matching Markets in Power Grid: Concepts and Solution using Deep Reinforcement Learning
- Learning Robust and Adaptive Real-World Continuous Control Using Simulation and Transfer Learning
- Improving On-policy Learning with Statistical Reward Accumulation