6 papers
PEARL: Training Socratic Tutors with Pedagogically Aligned Reinforcement Learning
Qikai Chang, Zhenrong Zhang, Linbo Chen +4
Large Language Models (LLMs) have shown promise as educational tutors, yet effective tutoring requires more than solving problems: it must provide progressive Socratic guidance and…
THOR: Tool-Integrated Hierarchical Optimization via RL for Mathematical Reasoning
Qikai Chang, Zhenrong Zhang, Pengfei Hu +6
Large Language Models (LLMs) have made remarkable progress in mathematical reasoning, but still continue to struggle with high-precision tasks like numerical computation and formal…
Step Potential Advantage Estimation: Harnessing Intermediate Confidence and Correctness for Efficient Mathematical Reasoning
Fei Wu, Zhenrong Zhang, Qikai Chang +3
Reinforcement Learning with Verifiable Rewards (RLVR) elicits long chain-of-thought reasoning in large language models (LLMs), but outcome-based rewards lead to coarse-grained adva…
PRM-BAS: Enhancing Multimodal Reasoning through PRM-guided Beam Annealing Search
Pengfei Hu, Zhenrong Zhang, Qikai Chang +8
Recent work increasingly focuses on improving the reasoning capabilities of Multimodal Large Language Models (MLLMs). Among existing methods, Process Reward Models (PRMs) stand out…
RFL: Simplifying Chemical Structure Recognition with Ring-Free Language
Qikai Chang, Mingjun Chen, Changpeng Pi +6
The primary objective of Optical Chemical Structure Recognition is to identify chemical structure images into corresponding markup sequences. However, the complex two-dimensional s…
Skeleton and Font Generation Network for Zero-shot Chinese Character Generation
Mobai Xue, Jun Du, Zhenrong Zhang +5
Automatic font generation remains a challenging research issue, primarily due to the vast number of Chinese characters, each with unique and intricate structures. Our investigation…