#mathematical reasoning
12 papers · 1 filter
SVR: Self-Verifying Refinement via Joint Verdict-Confidence Reinforcement Learning for Adaptive Test-Time Compute
Hongyu Chen, Liang Lin, Guangrun Wang
The paper proposes Self‑Verifying Refinement (SVR), a reinforcement‑learning framework that lets language models decide when to stop refining answers by using their own correctness…
LEEPS: Latent-Guided Explore-Exploit Prompt Sampling for Efficient RLVR in Large Language Models
Shuang Liang, Haoyang Zhou, Yifan Gong +2
The paper introduces LEEPS, a latent-guided explore‑exploit prompt sampler that selects prompts before rollout to reduce wasted generation budget and improve reinforcement learning…
LoRA Scaffolded Policy Optimization (LSPO): A Sampling-Time Low-Rank Scaffold for Recovering Reinforcement-Learning Gradient on Zero-Reward Cliff Prompts
Ken Ding
The paper introduces LoRA Scaffolded Policy Optimization (LSPO), a sampling-time low-rank adapter method that recovers gradient information for reinforcement learning on difficult…
Albilich: Steerable Proof-State Orchestration for LLM-Based Mathematical Research with CAS Integration
Ting Gong, Michael Ruofan Zeng, Yong Yang
Albilich is an open‑source agentic framework that lets large language models conduct long‑horizon mathematical research by integrating computer algebra systems, literature retrieva…
ReDiPPO: Reference-Guided Value Calibration and Discrepancy-Aware Token Reweighting for Mathematical Reasoning
Zhenrong Zhang, Fei Wu, Jun Du +2
The paper presents ReDiPPO, a PPO-based reinforcement learning framework that leverages reference answers to guide value estimation and reweights token-level advantages based on di…
ReCo: Reweighting GRPO Against Distributional Concentration
Junoh Park, Junseo Hwang, Wonguk Cho +1
The paper introduces ReCo, a reweighting technique for Group Relative Policy Optimization that mitigates the method’s tendency to focus on high‑probability responses, thereby impro…