4 papers
Soft Deterministic Policy Gradient with Gaussian Smoothing
Hyunjun Na, Donghwan Lee
Deterministic policy gradient (DPG) is widely utilized for continuous control; however, it inherently relies on the differentiability of the critic with respect to the action durin…
R-GTD: A Geometric Analysis of Gradient Temporal-Difference Learning in Singular Regimes
Hyunjun Na, Donghwan Lee
Gradient temporal-difference (GTD) learning algorithms are widely used for off-policy policy evaluation with function approximation. However, existing convergence analyses rely on…
Finite-Time Analysis of Simultaneous Double Q-learning
Hyunjun Na, Donghwan Lee
-learning is one of the most fundamental reinforcement learning (RL) algorithms. Despite its widespread success in various applications, it is prone to overestimation bias in th…
A Smooth Polynomial Lyapunov Certificate for Convergence of Q-Learning and Its Smooth Variants
Donghwan Lee, Hyunjun Na
Classical convergence analyses of Q-learning rely on the -norm contraction of Bellman operators, and existing ordinary differential equation (ODE) arguments often use the n…