2 papers
cs.LG2026
Finite-Time Accuracy of Temporal-Difference Learning Under Schur-Stable Recursions
Donghwan Lee, Do Wan Kim
Temporal difference (TD) learning is a cornerstone reinforcement learning (RL) method for policy evaluation, where the goal is to estimate the value function of a Markov decision p…
cs.LG2025
Regularized Q-learning
Han-Dong Lim, Donghwan Lee
Q-learning is widely used algorithm in reinforcement learning community. Under the lookup table setting, its convergence is well established. However, its behavior is known to be u…