Showing cs.AIShow all
2 papers · 1 filter
cs.AI2026
Don't Peek at the Answer: Outcome-Masked Group Relative Policy Optimization for Label-Free RLVR
Yongshi Ye, Liang Zhang, Yidong Chen +2
Reinforcement Learning with Verifiable Rewards (RLVR) improves LLM reasoning but typically relies on ground-truth (GT) answers, limiting scalability. Voting-based label-free RLVR r…
cs.AI2026
JURY-RL: Votes Propose, Proofs Dispose for Label-Free RLVR
Xinjie Chen, Biao Fu, Jing Wu +4
Reinforcement learning with verifiable rewards (RLVR) enhances the reasoning of large language models (LLMs), but standard RLVR often depends on human-annotated answers or carefull…