2 papers
cs.LG2024
Learning Infinite-Horizon Average-Reward Linear Mixture MDPs of Bounded Span
Woojin Chae, Kihyuk Hong, Yufan Zhang +2
This paper proposes a computationally tractable algorithm for learning infinite-horizon average-reward linear mixture Markov decision processes (MDPs) under the Bellman optimality…
cs.LG2024
Provably Efficient Infinite-Horizon Average-Reward Reinforcement Learning with Linear Function Approximation
Woojin Chae, Dabeen Lee
This paper proposes a computationally tractable algorithm for learning infinite-horizon average-reward linear Markov decision processes (MDPs) and linear mixture MDPs under the Bel…