2 papers
cs.LG2025
Offline RL with Smooth OOD Generalization in Convex Hull and its Neighborhood
Qingmao Yao, Zhichao Lei, Tianyuan Chen +5
Offline Reinforcement Learning (RL) struggles with distributional shifts, leading to the -value overestimation for out-of-distribution (OOD) actions. Existing methods address th…
cs.LG2024
Learning Diverse Policies with Soft Self-Generated Guidance
Guojian Wang, Faguo Wu, Xiao Zhang +1
Reinforcement learning (RL) with sparse and deceptive rewards is challenging because non-zero rewards are rarely obtained. Hence, the gradient calculated by the agent can be stocha…