activity
20212025
most citedEstimation Error Correction in Deep Reinforcement Learning for Deterministic Actor-Critic Methods

15 citations · 16 across the 8 of their papers we have counts for

collaborators

9 papers

cs.LG2025

Enhancing Deep Deterministic Policy Gradients on Continuous Control Tasks with Decoupled Prioritized Experience Replay

Mehmet Efe Lorasdagi, Dogan Can Cicek, Furkan Burak Mutlu +1

Background: Deep Deterministic Policy Gradient-based reinforcement learning algorithms utilize Actor-Critic architectures, where both networks are typically trained using identical…

cs.LG2024

CUER: Corrected Uniform Experience Replay for Off-Policy Continuous Deep Reinforcement Learning Algorithms

Arda Sarp Yenicesu, Furkan B. Mutlu, Suleyman S. Kozat +1

The utilization of the experience replay mechanism enables agents to effectively leverage their experiences on several occasions. In previous studies, the sampling probability of t…

cs.LG2022★ 1 cited

Actor Prioritized Experience Replay

Baturay Saglam, Furkan B. Mutlu, Dogan C. Cicek +1

A widely-studied deep reinforcement learning (RL) technique known as Prioritized Experience Replay (PER) allows agents to learn from transitions sampled with non-uniform probabilit…

cs.LG2022

Mitigating Off-Policy Bias in Actor-Critic Methods with One-Step Q-learning: A Novel Correction Approach

Baturay Saglam, Dogan C. Cicek, Furkan B. Mutlu +1

Compared to on-policy counterparts, off-policy model-free deep reinforcement learning can improve data efficiency by repeatedly using the previously gathered data. However, off-pol…

cs.LG2022

Safe and Robust Experience Sharing for Deterministic Policy Gradient Algorithms

Baturay Saglam, Dogan C. Cicek, Furkan B. Mutlu +1

Learning in high dimensional continuous tasks is challenging, mainly when the experience replay memory is very limited. We introduce a simple yet effective experience sharing mecha…

cs.LG2021

AWD3: Dynamic Reduction of the Estimation Bias

Dogan C. Cicek, Enes Duran, Baturay Saglam +3

Value-based deep Reinforcement Learning (RL) algorithms suffer from the estimation bias primarily caused by function approximation and temporal difference (TD) learning. This probl…