3 papers
cs.CL2025
PATS: Process-Level Adaptive Thinking Mode Switching
Yi Wang, Junxiao Liu, Shimao Zhang +2
Current large-language models (LLMs) typically adopt a fixed reasoning strategy, either simple or complex, for all questions, regardless of their difficulty. This neglect of variat…
cs.CL2025
R-PRM: Reasoning-Driven Process Reward Modeling
Shuaijie She, Junxiao Liu, Yifeng Liu +3
Large language models (LLMs) inevitably make mistakes when performing step-by-step mathematical reasoning. Process Reward Models (PRMs) have emerged as a promising solution by eval…
cs.CL2025
Process-based Self-Rewarding Language Models
Shimao Zhang, Xiao Liu, Xin Zhang +4
Large Language Models have demonstrated outstanding performance across various downstream tasks and have been widely applied in multiple scenarios. Human-annotated preference data…