7 papers
CriPO: Enhancing Rubric-based RL via Self-Distillation
Mingxuan Xia, Yuhang Yang, Chao Ye +7
Rubric-based RL has recently shown promise in improving LLMs on open-ended tasks. A widely recognized limitation of rubric-based RL is limited exploration: criteria that no rollout…
OPRD: On-Policy Representation Distillation
Shenzhi Yang, Guangcheng Zhu, Bowen Song +8
On-policy distillation (OPD) supervises the student exclusively in the output space by matching next-token distributions. This paradigm suffers from two limitations: (i) a high-var…
GeoMin: Data-Efficient Semi-Supervised RLVR via Geometric Distribution Modeling
Guangcheng Zhu, Shenzhi Yang, Haobo Wang +9
Reinforcement learning with verifiable rewards (RLVR) significantly advances LLM reasoning, yet it faces a dilemma: standard supervised scaling is throttled by high annotation cost…
Smart Picks in the Dark: Towards Efficient RLVR for Reasoning via Tracing Metacognitive Pivots
Guangcheng Zhu, Shenzhi Yang, Haobo Wang +7
Reinforcement learning with verifiable rewards (RLVR) has greatly advanced large reasoning models (LRMs), but it requires timely training on a huge fully-annotated dataset. To this…
Can LLMs Learn to Reason Robustly under Noisy Supervision?
Shenzhi Yang, Guangcheng Zhu, Bowen Song +7
Reinforcement Learning with Verifiable Rewards (RLVR) effectively trains reasoning models that rely on abundant perfect labels, but its vulnerability to unavoidable noisy labels du…
TraPO: A Semi-Supervised Reinforcement Learning Framework for Boosting LLM Reasoning
Shenzhi Yang, Guangcheng Zhu, Xing Zheng +7
Reinforcement learning with verifiable rewards (RLVR) has proven effective in training large reasoning models (LRMs) by leveraging answer-verifiable signals to guide policy optimiz…