4 papers
Your Language Model is Its Own Critic: Reinforcement Learning with Value Estimation from Actor's Internal States
Yunho Choi, Jongwon Lim, Woojin Ahn +3
Reinforcement learning with verifiable rewards (RLVR) for Large Reasoning Models hinges on baseline estimation for variance reduction, but existing approaches pay a heavy price: PP…
KL for a KL: On-Policy Distillation with Control Variate Baseline
Minjae Oh, Sangjun Song, Gyubin Choi +2
On-Policy Distillation (OPD) has emerged as a dominant post-training paradigm for large language models, especially for reasoning domains. However, OPD remains unstable in practice…
Future Policy Approximation for Offline Reinforcement Learning in LLM Reasoning
Minjae Oh, Yunho Choi, Dongmin Choi +1
Reinforcement learning (RL) has emerged as a key driver of post-training for complex reasoning in large language models (LLMs), yet online RL introduces substantial instability and…
Shallow-Ï: Knowledge Distillation for Flow-based VLAs
Boseong Jeon, Yunho Choi, Taehan Kim
The growing demand for real-time robotic deployment necessitates fast and on-device inference for vision-language-action (VLA) models. Within the VLA literature, efficiency has bee…