3 papers
cs.LG2025
Reward Models in Deep Reinforcement Learning: A Survey
Rui Yu, Shenghua Wan, Yucen Wang +4
In reinforcement learning (RL), agents continually interact with the environment and use the feedback to refine their behavior. To guide policy optimization, reward models are intr…
cs.LG2024
Hindsight Preference Learning for Offline Preference-based Reinforcement Learning
Chen-Xiao Gao, Shengjun Fang, Chenjun Xiao +2
Offline preference-based reinforcement learning (RL), which focuses on optimizing policies using human preferences between pairs of trajectory segments selected from an offline dat…
cs.LG2024
Disentangling Policy from Offline Task Representation Learning via Adversarial Data Augmentation
Chengxing Jia, Fuxiang Zhang, Yi-Chen Li +5
Offline meta-reinforcement learning (OMRL) proficiently allows an agent to tackle novel tasks while solely relying on a static dataset. For precise and efficient task identificatio…