3 papers
cs.LG2025
Finite-Time Bounds for Average-Reward Fitted Q-Iteration
Jongmin Lee, Ernest K. Ryu
Although there is an extensive body of work characterizing the sample complexity of discounted-return offline RL with function approximations, prior work on the average-reward sett…
cs.LG2025
Why Policy Gradient Algorithms Work for Undiscounted Total-Reward MDPs
Jongmin Lee, Ernest K. Ryu
The classical policy gradient method is the theoretical and conceptual foundation of modern policy-based reinforcement learning (RL) algorithms. Most rigorous analyses of such meth…
math.OC2025
Optimal Non-Asymptotic Rates of Value Iteration for Average-Reward Markov Decision Processes
Jongmin Lee, Ernest K. Ryu
While there is an extensive body of research on the analysis of Value Iteration (VI) for discounted cumulative-reward MDPs, prior work on analyzing VI for (undiscounted) average-re…