Showing cs.LGShow all
2 papers · 1 filter
cs.LG2026
Your Language Model is Its Own Critic: Reinforcement Learning with Value Estimation from Actor's Internal States
Yunho Choi, Jongwon Lim, Woojin Ahn +3
Reinforcement learning with verifiable rewards (RLVR) for Large Reasoning Models hinges on baseline estimation for variance reduction, but existing approaches pay a heavy price: PP…
cs.LG2026
KL for a KL: On-Policy Distillation with Control Variate Baseline
Minjae Oh, Sangjun Song, Gyubin Choi +2
On-Policy Distillation (OPD) has emerged as a dominant post-training paradigm for large language models, especially for reasoning domains. However, OPD remains unstable in practice…