5 papers
Safe-Support Q-Learning: Learning without Unsafe Exploration
Yeeun Lim, Narim Jeong, Donghwan Lee
Ensuring safety during reinforcement learning (RL) training is critical in real-world applications where unsafe exploration can lead to devastating outcomes. While most safe RL met…
Finite-Time Analysis of Q-Value Iteration for General-Sum Stackelberg Games
Narim Jeong, Donghwan Lee
Reinforcement learning has been successful both empirically and theoretically in single-agent settings, but extending these results to multi-agent reinforcement learning in general…
A Smooth Polynomial Lyapunov Certificate for Convergence of Q-Learning and Its Smooth Variants
Donghwan Lee, Hyunjun Na
Classical convergence analyses of Q-learning rely on the -norm contraction of Bellman operators, and existing ordinary differential equation (ODE) arguments often use the n…
Finite-Time Error Analysis of Soft Q-Learning: Switching System Approach
Narim Jeong, Donghwan Lee
Soft Q-learning is a variation of Q-learning designed to solve entropy regularized Markov decision problems where an agent aims to maximize the entropy regularized value function.…
Analysis of Off-Policy Multi-Step TD-Learning with Linear Function Approximation
Donghwan Lee
This paper analyzes multi-step TD-learning algorithms within the `deadly triad' scenario, characterized by linear function approximation, off-policy learning, and bootstrapping. In…