2 papers
cs.LG2024
Learning Infinite-Horizon Average-Reward Linear Mixture MDPs of Bounded Span
Woojin Chae, Kihyuk Hong, Yufan Zhang +2
This paper proposes a computationally tractable algorithm for learning infinite-horizon average-reward linear mixture Markov decision processes (MDPs) under the Bellman optimality…
stat.ML2024
Reinforcement Learning for Infinite-Horizon Average-Reward Linear MDPs via Approximation by Discounted-Reward MDPs
Kihyuk Hong, Woojin Chae, Yufan Zhang +2
We study the problem of infinite-horizon average-reward reinforcement learning with linear Markov decision processes (MDPs). The associated Bellman operator of the problem not bein…