2 papers
cs.LG2026
Multi-step Proximal Policy Improvement in Offline Reinforcement Learning
Soohyun Choi, Seonvin Cho, Songnam Hong
Offline reinforcement learning (RL) must reconcile two competing requirements: policy updates should stay near dataset-supported actions to keep value estimates reliable, yet meani…
cs.LG2026
PathBridger: Subgoal Bridges for Offline Goal-Conditioned Reinforcement Learning
Soohyun Choi, Seonvin Cho, Songnam Hong
Offline goal-conditioned reinforcement learning (GCRL) aims to learn policies for reaching diverse goals entirely from fixed trajectory data. Long-horizon offline GCRL remains chal…