3 papers
cs.LG2025
Semi-pessimistic Reinforcement Learning
Jin Zhu, Xin Zhou, Jiaang Yao +5
Offline reinforcement learning (RL) aims to learn an optimal policy from pre-collected data. However, it faces challenges of distributional shift, where the learned policy may enco…
cs.LG2023
Robust Offline Reinforcement learning with Heavy-Tailed Rewards
Jin Zhu, Runzhe Wan, Zhengling Qi +2
This paper endeavors to augment the robustness of offline reinforcement learning (RL) in scenarios laden with heavy-tailed rewards, a prevalent circumstance in real-world applicati…
stat.ML2022
An Instrumental Variable Approach to Confounded Off-Policy Evaluation
Yang Xu, Jin Zhu, Chengchun Shi +2
Off-policy evaluation (OPE) is a method for estimating the return of a target policy using some pre-collected observational data generated by a potentially different behavior polic…