2 papers
cs.AI2026
JURY-RL: Votes Propose, Proofs Dispose for Label-Free RLVR
Xinjie Chen, Biao Fu, Jing Wu +4
Reinforcement learning with verifiable rewards (RLVR) enhances the reasoning of large language models (LLMs), but standard RLVR often depends on human-annotated answers or carefull…
cs.LG2025
From Data-Centric to Sample-Centric: Enhancing LLM Reasoning via Progressive Optimization
Xinjie Chen, Minpeng Liao, Guoxin Chen +4
Reinforcement learning with verifiable rewards (RLVR) has recently advanced the reasoning capabilities of large language models (LLMs). While prior work has emphasized algorithmic…