4 papers
Latent Representation Alignment for Offline Goal-Conditioned Reinforcement Learning
Hyungkyu Kang, Byeongchan Kim, Min-hwan Oh
Offline goal-conditioned reinforcement learning (GCRL) provides a practical framework for obtaining goal-reaching policies from fixed datasets. However, learning a reliable goal-co…
Multi-Step Likelihood-Ratio Correction for Reinforcement Learning with Verifiable Rewards
Deokgyu Yoon, Hyungkyu Kang, Joongkyu Lee +4
Reinforcement learning with verifiable rewards (RLVR) plays a pivotal role in improving the reasoning ability of large language models. However, widely used PPO surrogate objective…
Peng's Q() for Conservative Value Estimation in Offline Reinforcement Learning
Byeongchan Kim, Min-hwan Oh
We propose a model-free offline multi-step reinforcement learning (RL) algorithm, Conservative Peng's Q() (CPQL). Our algorithm adapts the Peng's Q() (PQL) operator for con…
RelFlexformer: Efficient Attention 3D-Transformers for Integrable Relative Positional Encodings
Byeongchan Kim, Arijit Sehanobish, Avinava Dubey +2
We present a new class of efficient attention mechanisms applying universal 3D Relative Positional Encoding (RPE) methods given by arbitrary integrable modulation functions . Th…