collaborators

6 papers

cs.LG2026

PEARL: Training Socratic Tutors with Pedagogically Aligned Reinforcement Learning

Qikai Chang, Zhenrong Zhang, Linbo Chen +4

Large Language Models (LLMs) have shown promise as educational tutors, yet effective tutoring requires more than solving problems: it must provide progressive Socratic guidance and…

cs.AI2026

THOR: Tool-Integrated Hierarchical Optimization via RL for Mathematical Reasoning

Qikai Chang, Zhenrong Zhang, Pengfei Hu +6

Large Language Models (LLMs) have made remarkable progress in mathematical reasoning, but still continue to struggle with high-precision tasks like numerical computation and formal…

cs.CL2026

Step Potential Advantage Estimation: Harnessing Intermediate Confidence and Correctness for Efficient Mathematical Reasoning

Fei Wu, Zhenrong Zhang, Qikai Chang +3

Reinforcement Learning with Verifiable Rewards (RLVR) elicits long chain-of-thought reasoning in large language models (LLMs), but outcome-based rewards lead to coarse-grained adva…

cs.MM2025

PRM-BAS: Enhancing Multimodal Reasoning through PRM-guided Beam Annealing Search

Pengfei Hu, Zhenrong Zhang, Qikai Chang +8

Recent work increasingly focuses on improving the reasoning capabilities of Multimodal Large Language Models (MLLMs). Among existing methods, Process Reward Models (PRMs) stand out…

cs.CV2025

RFL: Simplifying Chemical Structure Recognition with Ring-Free Language

Qikai Chang, Mingjun Chen, Changpeng Pi +6

The primary objective of Optical Chemical Structure Recognition is to identify chemical structure images into corresponding markup sequences. However, the complex two-dimensional s…

cs.CV2025

Skeleton and Font Generation Network for Zero-shot Chinese Character Generation

Mobai Xue, Jun Du, Zhenrong Zhang +5

Automatic font generation remains a challenging research issue, primarily due to the vast number of Chinese characters, each with unique and intricate structures. Our investigation…