large language models 1mathematical reasoning 1policy optimization 1reinforcement learning 1value estimation 1
From the 1 of 12 linked papers with an AI index.
Showing cs.CLShow all
3 papers · 1 filter
cs.CL2026
Step Potential Advantage Estimation: Harnessing Intermediate Confidence and Correctness for Efficient Mathematical Reasoning
Fei Wu, Zhenrong Zhang, Qikai Chang +3
Reinforcement Learning with Verifiable Rewards (RLVR) elicits long chain-of-thought reasoning in large language models (LLMs), but outcome-based rewards lead to coarse-grained adva…
cs.CL2025
Enhancing the Geometric Problem-Solving Ability of Multimodal LLMs via Symbolic-Neural Integration
Yicheng Pan, Zhenrong Zhang, Pengfei Hu +6
Recent advances in Multimodal Large Language Models (MLLMs) have achieved remarkable progress in general domains and demonstrated promise in multimodal mathematical reasoning. Howe…
cs.CL2025
DocMamba: Efficient Document Pre-training with State Space Model
Pengfei Hu, Zhenrong Zhang, Jiefeng Ma +3
In recent years, visually-rich document understanding has attracted increasing attention. Transformer-based pre-trained models have become the mainstream approach, yielding signifi…