4 papers
An Optimistic Algorithm for online CMDPS with Anytime Adversarial Constraints
Jiahui Zhu, Kihyun Yu, Dabeen Lee +2
Online safe reinforcement learning (RL) plays a key role in dynamic environments, with applications in autonomous driving, robotics, and cybersecurity. The objective is to learn op…
Learning Infinite-Horizon Average-Reward Linear Mixture MDPs of Bounded Span
Woojin Chae, Kihyuk Hong, Yufan Zhang +2
This paper proposes a computationally tractable algorithm for learning infinite-horizon average-reward linear mixture Markov decision processes (MDPs) under the Bellman optimality…
Improved Regret Bound for Safe Reinforcement Learning via Tighter Cost Pessimism and Reward Optimism
Kihyun Yu, Duksang Lee, William Overman +1
This paper studies the safe reinforcement learning problem formulated as an episodic finite-horizon tabular constrained Markov decision process with an unknown transition kernel an…
Infinite-Horizon Reinforcement Learning with Multinomial Logistic Function Approximation
Jaehyun Park, Junyeop Kwon, Dabeen Lee
We study model-based reinforcement learning with non-linear function approximation where the transition function of the underlying Markov decision process (MDP) is given by a multi…