2 papers
cs.LG2026
Frozen Policy Iteration: Computationally Efficient RL under Linear Realizability for Deterministic Dynamics
Yijing Ke, Zihan Zhang, Ruosong Wang
We study computationally and statistically efficient reinforcement learning under the linear realizability assumption, where any policy's -function is linear in a given s…
cs.LG2024
Uniform Last-Iterate Guarantee for Bandits and Reinforcement Learning
Junyan Liu, Yunfan Li, Ruosong Wang +1
Existing metrics for reinforcement learning (RL) such as regret, PAC bounds, or uniform-PAC (Dann et al., 2017), typically evaluate the cumulative performance, while allowing the a…