4 papers
Policy Gradient Algorithms in Average-Reward Multichain MDPs
Jongmin Lee, Ernest K. Ryu
While there is an extensive body of research analyzing policy gradient methods for discounted cumulative-reward MDPs, prior work on policy gradient methods for average-reward MDPs…
Why Policy Gradient Algorithms Work for Undiscounted Total-Reward MDPs
Jongmin Lee, Ernest K. Ryu
The classical policy gradient method is the theoretical and conceptual foundation of modern policy-based reinforcement learning (RL) algorithms. Most rigorous analyses of such meth…
Finite-Time Bounds for Average-Reward Fitted Q-Iteration
Jongmin Lee, Ernest K. Ryu
Although there is an extensive body of work characterizing the sample complexity of discounted-return offline RL with function approximations, prior work on the average-reward sett…
Deflated Dynamics Value Iteration
Jongmin Lee, Amin Rakhsha, Ernest K. Ryu +1
The Value Iteration (VI) algorithm is an iterative procedure to compute the value function of a Markov decision process, and is the basis of many reinforcement learning (RL) algori…