most citedRethinking Sample Polarity in Reinforcement Learning with Verifiable Rewards

1 citations · 1 across the 6 of their papers we have counts for

collaborators

7 papers

cs.CL20251 cited

Rethinking Sample Polarity in Reinforcement Learning with Verifiable Rewards

Xinyu Tang, Yuliang Zhan, Zhixun Li +5

Large reasoning models (LRMs) are typically trained using reinforcement learning with verifiable reward (RLVR) to enhance their reasoning abilities. In this paradigm, policies are…

cs.CL2025

Every Step Evolves: Scaling Reinforcement Learning for Trillion-Scale Thinking Model

Ling Team, Anqi Shen, Baihui Li +101

We present Ring-1T, the first open-source, state-of-the-art thinking model with a trillion-scale parameter. It features 1 trillion total parameters and activates approximately 50 b…

cs.LG2025

Towards High Data Efficiency in Reinforcement Learning with Verifiable Reward

Xinyu Tang, Zhenduo Zhang, Yurou Liu +4

Recent advances in large reasoning models have leveraged reinforcement learning with verifiable rewards (RLVR) to improve reasoning capabilities. However, scaling these methods typ…

cs.CL2025

Ring-lite: Scalable Reasoning via C3PO-Stabilized Reinforcement Learning for LLMs

Ling Team, Bin Hu, Cai Chen +43

We present Ring-lite, a Mixture-of-Experts (MoE)-based large language model optimized via reinforcement learning (RL) to achieve efficient and robust reasoning capabilities. Built…

cs.AI2025

SHARP: Synthesizing High-quality Aligned Reasoning Problems for Large Reasoning Models Reinforcement Learning

Xiong Jun Wu, Zhenduo Zhang, ZuJie Wen +11

Training large reasoning models (LRMs) with reinforcement learning in STEM domains is hindered by the scarcity of high-quality, diverse, and verifiable problem sets. Existing synth…

cs.AI2025

Incentivizing Dual Process Thinking for Efficient Large Language Model Reasoning

Xiaoxue Cheng, Junyi Li, Zhenduo Zhang +4

Large reasoning models (LRMs) have demonstrated strong performance on complex reasoning tasks, but often suffer from overthinking, generating redundant content regardless of task d…