works on

From the 1 of 11 linked papers with an AI index.

collaborators

11 papers

cs.AI2026

ReDiPPO: Reference-Guided Value Calibration and Discrepancy-Aware Token Reweighting for Mathematical Reasoning

Zhenrong Zhang, Fei Wu, Jun Du +2

The paper presents ReDiPPO, a PPO-based reinforcement learning framework that leverages reference answers to guide value estimation and reweights token-level advantages based on di…

cs.AI2026

EquivPruner: Boosting Efficiency and Quality in LLM-Based Search via Action Pruning

Jiawei Liu, Qisi Chen, Jianshu Zhang +2

Large Language Models (LLMs) excel at complex reasoning through search algorithms, yet current strategies often suffer from massive token consumption due to redundant exploration o…

cs.LG2026

PEARL: Training Socratic Tutors with Pedagogically Aligned Reinforcement Learning

Qikai Chang, Zhenrong Zhang, Linbo Chen +4

Large Language Models (LLMs) have shown promise as educational tutors, yet effective tutoring requires more than solving problems: it must provide progressive Socratic guidance and…

cs.AI2026

THOR: Tool-Integrated Hierarchical Optimization via RL for Mathematical Reasoning

Qikai Chang, Zhenrong Zhang, Pengfei Hu +6

Large Language Models (LLMs) have made remarkable progress in mathematical reasoning, but still continue to struggle with high-precision tasks like numerical computation and formal…

cs.CL2026

Step Potential Advantage Estimation: Harnessing Intermediate Confidence and Correctness for Efficient Mathematical Reasoning

Fei Wu, Zhenrong Zhang, Qikai Chang +3

Reinforcement Learning with Verifiable Rewards (RLVR) elicits long chain-of-thought reasoning in large language models (LLMs), but outcome-based rewards lead to coarse-grained adva…

cs.CV2025

See then Tell: Enhancing Key Information Extraction with Vision Grounding

Shuhang Liu, Zhenrong Zhang, Pengfei Hu +5

In the digital era, the ability to understand visually rich documents that integrate text, complex layouts, and imagery is critical. Traditional Key Information Extraction (KIE) me…