7 papers · 1 filter
Be My Tutor: On-Policy Co-Distillation for Mutual LLM Improvement via Peer Feedback
Woohyeon Byeon, Jiwon Jeon, Jeonghye Kim +1
We study multi-domain LLM training in which two models, each stronger in a different domain, co-evolve by tutoring each other through on-policy feedback. Unlike one-way distillatio…
Rebellious Student: Reversing Teacher Signals for Reasoning Exploration with Self-Distilled RLVR
Jeonghye Kim, Jiwon Jeon, Dongsheng Li +1
Self-distillation has emerged as a powerful framework for post-training LLMs, where a teacher conditioned on extra information guides a student without it, both from the same model…
Exploratory Memory-Augmented LLM Agent via Hybrid On- and Off-Policy Optimization
Zeyuan Liu, Jeonghye Kim, Xufang Luo +2
Exploration remains the key bottleneck for large language model agents trained with reinforcement learning. While prior methods exploit pretrained knowledge, they fail in environme…
Penalizing Infeasible Actions and Reward Scaling in Reinforcement Learning with Offline Data
Jeonghye Kim, Yongjae Shin, Whiyoung Jung +5
Reinforcement learning with offline data suffers from Q-value extrapolation errors. To address this issue, we first demonstrate that linear extrapolation of the Q-function beyond t…
Adaptive -Aid for Conditional Supervised Learning in Offline Reinforcement Learning
Jeonghye Kim, Suyoung Lee, Woojun Kim +1
Offline reinforcement learning (RL) has progressed with return-conditioned supervised learning (RCSL), but its lack of stitching ability remains a limitation. We introduce -Aide…
LESSON: Learning to Integrate Exploration Strategies for Reinforcement Learning via an Option Framework
Woojun Kim, Jeonghye Kim, Youngchul Sung
In this paper, a unified framework for exploration in reinforcement learning (RL) is proposed based on an option-critic model. The proposed framework learns to integrate a set of d…