Showing cs.LGShow all
3 papers · 1 filter
cs.LG2026
Learning Policy from a Single Trajectory in Average-Reward Markov Decision Process
Jongmin Lee, Ernest K. Ryu, Vaneet Aggarwal
While there is an extensive body of work characterizing the sample complexity of discounted cumulative-reward MDPs, finite sample analyses for average-reward MDPs have been limited…
cs.LG2025
Finite-Time Bounds for Average-Reward Fitted Q-Iteration
Jongmin Lee, Ernest K. Ryu
Although there is an extensive body of work characterizing the sample complexity of discounted-return offline RL with function approximations, prior work on the average-reward sett…
cs.LG2025
Why Policy Gradient Algorithms Work for Undiscounted Total-Reward MDPs
Jongmin Lee, Ernest K. Ryu
The classical policy gradient method is the theoretical and conceptual foundation of modern policy-based reinforcement learning (RL) algorithms. Most rigorous analyses of such meth…