1 paper · 1 filter
Byeongho Heo, Jaehui Hwang, Sangdoo Yun +1
On-policy distillation is an alternative post-training method in reinforcement learning that alleviates the constraints imposed by reward models by providing token-level supervisio…