1 paper · 1 filter
Xue Gong, Qi Yi, Ziyuan Nan +8
Training Large Language Models (LLMs) for reasoning tasks is increasingly driven by Reinforcement Learning with Verifiable Rewards (RLVR), where Proximal Policy Optimization (PPO)…