3 papers
cs.LG2025
Soft Adaptive Policy Optimization
Chang Gao, Chujie Zheng, Xiong-Hui Chen +7
Reinforcement learning (RL) plays an increasingly important role in enhancing the reasoning capabilities of large language models (LLMs), yet stable and performant policy optimizat…
cs.AI2025
StepFun-Prover Preview: Let's Think and Verify Step by Step
Shijie Shang, Ruosi Wan, Yue Peng +4
We present StepFun-Prover Preview, a large language model designed for formal theorem proving through tool-integrated reasoning. Using a reinforcement learning pipeline that incorp…
cs.CL2025
Beyond the 80/20 Rule: High-Entropy Minority Tokens Drive Effective Reinforcement Learning for LLM Reasoning
Shenzhi Wang, Le Yu, Chang Gao +15
Reinforcement Learning with Verifiable Rewards (RLVR) has emerged as a powerful approach to enhancing the reasoning capabilities of Large Language Models (LLMs), while its mechanis…