Revisiting Experience Replayable Conditions
arXiv:2402.10374 · doi:10.1007/s10489-024-05685-7
Abstract
Experience replay (ER) used in (deep) reinforcement learning is considered to be applicable only to off-policy algorithms. However, there have been some cases in which ER has been applied for on-policy algorithms, suggesting that off-policyness might be a sufficient condition for applying ER. This paper reconsiders more strict "experience replayable conditions" (ERC) and proposes the way of modifying the existing algorithms to satisfy ERC. In light of this, it is postulated that the instability of policy improvements represents a pivotal factor in ERC. The instability factors are revealed from the viewpoint of metric learning as i) repulsive forces from negative samples and ii) replays of inappropriate experiences. Accordingly, the corresponding stabilization tricks are derived. As a result, it is confirmed through numerical simulations that the proposed stabilization tricks make ER applicable to an advantage actor-critic, an on-policy algorithm. Moreover, its learning performance is comparable to that of a soft actor-critic, a state-of-the-art off-policy algorithm.
25 pages, 9 figures
References in corpus (7)
- dm_control: Software and Tasks for Continuous Control
- AdaTerm: Adaptive T-Distribution Estimated Robust Moments for Noise-Robust Stochastic Gradient Optimization
- Optimistic Reinforcement Learning by Forward Kullback-Leibler Divergence Optimization
- L2C2: Locally Lipschitz Continuous Constraint towards Stable and Smooth Reinforcement Learning
- Proximal Policy Optimization with Adaptive Threshold for Symmetric Relative Density Ratio
- Design of Restricted Normalizing Flow towards Arbitrary Stochastic Policy with Computational Efficiency
- Consolidated Adaptive T-soft Update for Deep Reinforcement Learning