large language models 1mathematical reasoning 1policy optimization 1reinforcement learning 1value estimation 1
From the 1 of 2 linked papers with an AI index.
2 papers
cs.AI2026
ReDiPPO: Reference-Guided Value Calibration and Discrepancy-Aware Token Reweighting for Mathematical Reasoning
Zhenrong Zhang, Fei Wu, Jun Du +2
The paper presents ReDiPPO, a PPO-based reinforcement learning framework that leverages reference answers to guide value estimation and reweights token-level advantages based on di…
cs.CL2025
Spark-Prover-X1: Formal Theorem Proving Through Diverse Data Training
Xinyuan Zhou, Yi Lei, Xiaoyu Zhou +7
Large Language Models (LLMs) have shown significant promise in automated theorem proving, yet progress is often constrained by the scarcity of diverse and high-quality formal langu…