A Deeper Look at Experience Replay
arXiv:1712.01275
Abstract
Recently experience replay is widely used in various deep reinforcement learning (RL) algorithms, in this paper we rethink the utility of experience replay. It introduces a new hyper-parameter, the memory buffer size, which needs carefully tuning. However unfortunately the importance of this new hyper-parameter has been underestimated in the community for a long time. In this paper we did a systematic empirical study of experience replay under various function representations. We showcase that a large replay buffer can significantly hurt the performance. Moreover, we propose a simple O(1) method to remedy the negative influence of a large replay buffer. We showcase its utility in both simple grid world and challenging domains like Atari games.
NIPS 2017 Deep Reinforcement Learning Symposium
References in corpus (3)
Cited by in corpus (44)
- Off-Policy Deep Reinforcement Learning without Exploration
- Multi-UAV Path Planning for Wireless Data Harvesting with Deep Reinforcement Learning
- Go-Explore: a New Approach for Hard-Exploration Problems
- A Theoretical Analysis of Deep Q-Learning
- Review: Deep Learning in Electron Microscopy
- Learning to Learn without Forgetting by Maximizing Transfer and Minimizing Interference
- An Optimistic Perspective on Offline Reinforcement Learning
- UAV Path Planning using Global and Local Map Information with Deep Reinforcement Learning
- DisCor: Corrective Feedback in Reinforcement Learning via Distribution Correction
- Recurrent Deterministic Policy Gradient Method for Bipedal Locomotion on Rough Terrain Challenge
- Prioritized Sequence Experience Replay
- Deep reinforcement learning in World-Earth system models to discover sustainable management strategies
- Physical Informed-Inspired Deep Reinforcement Learning Based Bi-Level Programming for Microgrid Scheduling
- Modeling Interactions of Autonomous Vehicles and Pedestrians with Deep Multi-Agent Reinforcement Learning for Collision Avoidance
- ARCHER: Aggressive Rewards to Counter bias in Hindsight Experience Replay
- Experience Replay Optimization
- Time Limits in Reinforcement Learning
- Simplifying Deep Reinforcement Learning via Self-Supervision
- Cloud-Edge Training Architecture for Sim-to-Real Deep Reinforcement Learning
- An Equivalence between Loss Functions and Non-Uniform Sampling in Experience Replay
- A Memory Efficient Deep Reinforcement Learning Approach For Snake Game Autonomous Agents
- Continual Reinforcement Learning with Multi-Timescale Replay
- Improving Performance in Reinforcement Learning by Breaking Generalization in Neural Networks
- ZPD Teaching Strategies for Deep Reinforcement Learning from Demonstrations
- Learning Sparse Representations Incrementally in Deep Reinforcement Learning
- Interactive Visualization for Debugging RL
- Toolpath design for additive manufacturing using deep reinforcement learning
- Rank the Episodes: A Simple Approach for Exploration in Procedurally-Generated Environments
- Off-Policy Risk-Sensitive Reinforcement Learning Based Constrained Robust Optimal Control
- Experience Augmentation: Boosting and Accelerating Off-Policy Multi-Agent Reinforcement Learning
- Imaginary Hindsight Experience Replay: Curious Model-based Learning for Sparse Reward Tasks
- Shared learning of powertrain control policies for vehicle fleets
- Learning to Sample with Local and Global Contexts in Experience Replay Buffer
- Improving Experience Replay through Modeling of Similar Transitions' Sets
- Deep Reinforcement Learning Based Robot Arm Manipulation with Efficient Training Data through Simulation
- Recomposing the Reinforcement Learning Building Blocks with Hypernetworks
- Discrete-to-Deep Supervised Policy Learning
- Learning Wildfire Model from Incomplete State Observations
- Self-Improving Semantic Perception for Indoor Localisation
- Lucid Dreaming for Experience Replay: Refreshing Past States with the Current Policy
- DCUR: Data Curriculum for Teaching via Samples with Reinforcement Learning
- FiDi-RL: Incorporating Deep Reinforcement Learning with Finite-Difference Policy Search for Efficient Learning of Continuous Control
- Adaptive Experience Selection for Policy Gradient
- Large Batch Experience Replay