2 papers
cs.LG2026
Persistent Teacher Anchoring for Tool-Using Agents
Hyun Bin Park, Kyungho Song, Sangmin Lee +1
Distillation is common in LLM post-training, where on-policy knowledge distillation (OPKD) uses student-generated trajectories to prepare the student for downstream RL. At each sta…
cs.LG2026
Headroom-Drift Replay: A Primitive for Principled Replay Control in GRPO
Hyun Bin Park, Du-Seong Chang
RL-based post-training for reasoning models is increasingly bottlenecked by repeated fresh rollout generation, particularly in agentic settings where environment interaction domina…