2 papers
cs.LG2026
SALT: When More Rollouts Don't Help in Group-Based Policy Optimization and How to Make Them Matter
Powei Chang, Jinpeng Zhang, Chaoqun Sun +6
Reinforcement learning with verifiable rewards (RLVR) often adopts GRPO-style group-relative updates, sampling multiple rollouts per prompt to construct normalized learning signals…
cs.AI2026
ScRPO: From Errors to Insights
Lianrui Li, Dakuan Lu, Jiawei Shao +1
We introduce Self-correction Relative Policy Optimization (ScRPO), a novel reinforcement learning framework designed to empower large language models with advanced mathematical rea…