15 citations · 15 across the 2 of their papers we have counts for
3 papers
AWD3: Dynamic Reduction of the Estimation Bias
Dogan C. Cicek, Enes Duran, Baturay Saglam +3
Value-based deep Reinforcement Learning (RL) algorithms suffer from the estimation bias primarily caused by function approximation and temporal difference (TD) learning. This probl…
Off-Policy Correction for Deep Deterministic Policy Gradient Algorithms via Batch Prioritized Experience Replay
Dogan C. Cicek, Enes Duran, Baturay Saglam +2
The experience replay mechanism allows agents to use the experiences multiple times. In prior works, the sampling probability of the transitions was adjusted according to their imp…
Estimation Error Correction in Deep Reinforcement Learning for Deterministic Actor-Critic Methods
Baturay Saglam, Enes Duran, Dogan C. Cicek +2
In value-based deep reinforcement learning methods, approximation of value functions induces overestimation bias and leads to suboptimal policies. We show that in deep actor-critic…