Showing cs.LGShow all
3 papers · 1 filter
cs.LG2026
Be My Tutor: On-Policy Co-Distillation for Mutual LLM Improvement via Peer Feedback
Woohyeon Byeon, Jiwon Jeon, Jeonghye Kim +1
We study multi-domain LLM training in which two models, each stronger in a different domain, co-evolve by tutoring each other through on-policy feedback. Unlike one-way distillatio…
cs.LG2026
Adaptive Action Chunking via Multi-Chunk Q Value Estimation
Yongjae Shin, Jongseong Chae, Seongmin Kim +2
Action chunking emerged as a pivotal technique in imitation learning, enabling policies to predict cohesive action sequences rather than single actions. Recently, this approach has…
cs.LG2025
Penalizing Infeasible Actions and Reward Scaling in Reinforcement Learning with Offline Data
Jeonghye Kim, Yongjae Shin, Whiyoung Jung +5
Reinforcement learning with offline data suffers from Q-value extrapolation errors. To address this issue, we first demonstrate that linear extrapolation of the Q-function beyond t…