1 paper · 1 filter
Haizhong Zheng, Yang Zhou, Brian R. Bartoldson +4
Reinforcement learning, such as PPO and GRPO, has powered recent breakthroughs in LLM reasoning. Scaling rollout to sample more prompts enables models to selectively use higher-qua…