3 papers
cs.LG2025
Preference-Guided Reinforcement Learning for Efficient Exploration
Guojian Wang, Jianxiang Liu, Xinyuan Li +4
In this paper, we investigate preference-based reinforcement learning (PbRL), which enables reinforcement learning (RL) agents to learn from human feedback. This is particularly va…
cs.LG2025
Offline RL with Smooth OOD Generalization in Convex Hull and its Neighborhood
Qingmao Yao, Zhichao Lei, Tianyuan Chen +5
Offline Reinforcement Learning (RL) struggles with distributional shifts, leading to the -value overestimation for out-of-distribution (OOD) actions. Existing methods address th…
cs.LG2024
Policy Optimization with Smooth Guidance Learned from State-Only Demonstrations
Guojian Wang, Faguo Wu, Xiao Zhang +1
The sparsity of reward feedback remains a challenging problem in online deep reinforcement learning (DRL). Previous approaches have utilized offline demonstrations to achieve impre…