#mathematical reasoning

topicmathematical reasoning

12 papers · 1 filter

cs.AI2026

SVR: Self-Verifying Refinement via Joint Verdict-Confidence Reinforcement Learning for Adaptive Test-Time Compute

Hongyu Chen, Liang Lin, Guangrun Wang

The paper proposes Self‑Verifying Refinement (SVR), a reinforcement‑learning framework that lets language models decide when to stop refining answers by using their own correctness…

cs.CL2026

LEEPS: Latent-Guided Explore-Exploit Prompt Sampling for Efficient RLVR in Large Language Models

Shuang Liang, Haoyang Zhou, Yifan Gong +2

The paper introduces LEEPS, a latent-guided explore‑exploit prompt sampler that selects prompts before rollout to reduce wasted generation budget and improve reinforcement learning…

cs.LG2026

LoRA Scaffolded Policy Optimization (LSPO): A Sampling-Time Low-Rank Scaffold for Recovering Reinforcement-Learning Gradient on Zero-Reward Cliff Prompts

Ken Ding

The paper introduces LoRA Scaffolded Policy Optimization (LSPO), a sampling-time low-rank adapter method that recovers gradient information for reinforcement learning on difficult…

cs.AI2026

Albilich: Steerable Proof-State Orchestration for LLM-Based Mathematical Research with CAS Integration

Ting Gong, Michael Ruofan Zeng, Yong Yang

Albilich is an open‑source agentic framework that lets large language models conduct long‑horizon mathematical research by integrating computer algebra systems, literature retrieva…

cs.AI2026

ReDiPPO: Reference-Guided Value Calibration and Discrepancy-Aware Token Reweighting for Mathematical Reasoning

Zhenrong Zhang, Fei Wu, Jun Du +2

The paper presents ReDiPPO, a PPO-based reinforcement learning framework that leverages reference answers to guide value estimation and reweights token-level advantages based on di…

cs.LG2026

ReCo: Reweighting GRPO Against Distributional Concentration

Junoh Park, Junseo Hwang, Wonguk Cho +1

The paper introduces ReCo, a reweighting technique for Group Relative Policy Optimization that mitigates the method’s tendency to focus on high‑probability responses, thereby impro…