1 paper
Haizhong Zheng, Jiawei Zhao, Beidi Chen
Reinforcement learning has been central to recent advances in large language model reasoning, but most algorithms rely on on-policy training that demands fresh rollouts at every up…