4 papers
Finite-Time Accuracy of Temporal-Difference Learning Under Schur-Stable Recursions
Donghwan Lee, Do Wan Kim
Temporal difference (TD) learning is a cornerstone reinforcement learning (RL) method for policy evaluation, where the goal is to estimate the value function of a Markov decision p…
Backstepping Temporal Difference Learning
Han-Dong Lim, Donghwan Lee
Off-policy learning ability is an important feature of reinforcement learning (RL) for practical applications. However, even one of the most elementary RL algorithms, temporal-diff…
Finite-Time Analysis of Temporal Difference Learning with Experience Replay
Han-Dong Lim, Donghwan Lee
Temporal-difference (TD) learning is widely regarded as one of the most popular algorithms in reinforcement learning (RL). Despite its widespread use, it has only been recently tha…
Continuous-Time Distributed Dynamic Programming for Networked Multi-Agent Markov Decision Processes
Donghwan Lee, Han-Dong Lim, Do Wan Kim
The main goal of this paper is to investigate continuous-time distributed dynamic programming (DP) algorithms for networked multi-agent Markov decision problems (MAMDPs). In our st…