4 papers
Adaptive Action Chunking via Multi-Chunk Q Value Estimation
Yongjae Shin, Jongseong Chae, Seongmin Kim +2
Action chunking emerged as a pivotal technique in imitation learning, enabling policies to predict cohesive action sequences rather than single actions. Recently, this approach has…
Flow Matching with Injected Noise for Offline-to-Online Reinforcement Learning
Yongjae Shin, Jongseong Chae, Jongeui Park +1
Generative models have recently demonstrated remarkable success across diverse domains, motivating their adoption as expressive policies in reinforcement learning (RL). While they…
Flow Actor-Critic for Offline Reinforcement Learning
Jongseong Chae, Jongeui Park, Yongjae Shin +3
The dataset distributions in offline reinforcement learning (RL) often exhibit complex and multi-modal distributions, necessitating expressive policies to capture such distribution…
Penalizing Infeasible Actions and Reward Scaling in Reinforcement Learning with Offline Data
Jeonghye Kim, Yongjae Shin, Whiyoung Jung +5
Reinforcement learning with offline data suffers from Q-value extrapolation errors. To address this issue, we first demonstrate that linear extrapolation of the Q-function beyond t…