most citedRethinking Sample Polarity in Reinforcement Learning with Verifiable Rewards

1 citations · 1 across the 6 of their papers we have counts for

collaborators

8 papers

cs.LG2026

When Sharpening Becomes Collapse: Sampling Bias and Semantic Coupling in RL with Verifiable Rewards

Mingyuan Fan, Weiguang Han, Daixin Wang +3

Reinforcement Learning with Verifiable Rewards (RLVR) is a central paradigm for turning large language models (LLMs) into reliable problem solvers, especially in logic-heavy domain…

cs.CL20251 cited

Rethinking Sample Polarity in Reinforcement Learning with Verifiable Rewards

Xinyu Tang, Yuliang Zhan, Zhixun Li +5

Large reasoning models (LRMs) are typically trained using reinforcement learning with verifiable reward (RLVR) to enhance their reasoning abilities. In this paradigm, policies are…

cs.LG2025

Towards High Data Efficiency in Reinforcement Learning with Verifiable Reward

Xinyu Tang, Zhenduo Zhang, Yurou Liu +4

Recent advances in large reasoning models have leveraged reinforcement learning with verifiable rewards (RLVR) to improve reasoning capabilities. However, scaling these methods typ…

cs.CL2025

Ring-lite: Scalable Reasoning via C3PO-Stabilized Reinforcement Learning for LLMs

Ling Team, Bin Hu, Cai Chen +43

We present Ring-lite, a Mixture-of-Experts (MoE)-based large language model optimized via reinforcement learning (RL) to achieve efficient and robust reasoning capabilities. Built…

cs.AI2025

SHARP: Synthesizing High-quality Aligned Reasoning Problems for Large Reasoning Models Reinforcement Learning

Xiong Jun Wu, Zhenduo Zhang, ZuJie Wen +11

Training large reasoning models (LRMs) with reinforcement learning in STEM domains is hindered by the scarcity of high-quality, diverse, and verifiable problem sets. Existing synth…

cs.CL2025

MASS: Mathematical Data Selection via Skill Graphs for Pretraining Large Language Models

Jiazheng Li, Lu Yu, Qing Cui +4

High-quality data plays a critical role in the pretraining and fine-tuning of large language models (LLMs), even determining their performance ceiling to some degree. Consequently,…