large language models 1mathematical reasoning 1policy optimization 1reinforcement learning 1value estimation 1
From the 1 of 12 linked papers with an AI index.
Showing cs.AIShow all
2 papers · 1 filter
cs.AI2026
ReDiPPO: Reference-Guided Value Calibration and Discrepancy-Aware Token Reweighting for Mathematical Reasoning
Zhenrong Zhang, Fei Wu, Jun Du +2
The paper presents ReDiPPO, a PPO-based reinforcement learning framework that leverages reference answers to guide value estimation and reweights token-level advantages based on di…
cs.AI2026
THOR: Tool-Integrated Hierarchical Optimization via RL for Mathematical Reasoning
Qikai Chang, Zhenrong Zhang, Pengfei Hu +6
Large Language Models (LLMs) have made remarkable progress in mathematical reasoning, but still continue to struggle with high-precision tasks like numerical computation and formal…