Showing cs.LGShow all
3 papers · 1 filter
cs.LG2026
Why Policy Gradient Algorithms Work for Undiscounted Total-Reward MDPs
Jongmin Lee, Ernest K. Ryu
The classical policy gradient method is the theoretical and conceptual foundation of modern policy-based reinforcement learning (RL) algorithms. Most rigorous analyses of such meth…
cs.LG2025
Finite-Time Bounds for Average-Reward Fitted Q-Iteration
Jongmin Lee, Ernest K. Ryu
Although there is an extensive body of work characterizing the sample complexity of discounted-return offline RL with function approximations, prior work on the average-reward sett…
cs.LG2025
Deflated Dynamics Value Iteration
Jongmin Lee, Amin Rakhsha, Ernest K. Ryu +1
The Value Iteration (VI) algorithm is an iterative procedure to compute the value function of a Markov decision process, and is the basis of many reinforcement learning (RL) algori…