6 papers
RP-OPSD: Reasoning-Pivot-Guided On-Policy Self-Distillation for Multilingual Reasoning Transfer
Xinye Wang, Junxiao Liu, Shujian Huang
Multilingual reasoning transfer is crucial for extending reasoning capabilities of large language models (LLMs) beyond high-resource languages. On-policy self-distillation (OPSD) a…
Efficient Multilingual Reasoning Transfer via Progressive Code-Switching
Zhijun Wang, Junxiao Liu, Hao Zhou +3
Large reasoning models (LRMs) have achieved strong reasoning capabilities in English, yet their performance degrades significantly when required to reason in other languages. A nat…
R3S: Refining and Recovering Reinforcement Signals for Multilingual Understanding and Reasoning
Junxiao Liu, Zhijun Wang, Yixiao Li +6
Large reasoning models often default to English reasoning when processing non-English questions, yet their performance drops substantially when reasoning in the question language.…
PATS: Process-Level Adaptive Thinking Mode Switching
Yi Wang, Junxiao Liu, Shimao Zhang +2
Current large-language models (LLMs) typically adopt a fixed reasoning strategy, either simple or complex, for all questions, regardless of their difficulty. This neglect of variat…
R-PRM: Reasoning-Driven Process Reward Modeling
Shuaijie She, Junxiao Liu, Yifeng Liu +3
Large language models (LLMs) inevitably make mistakes when performing step-by-step mathematical reasoning. Process Reward Models (PRMs) have emerged as a promising solution by eval…
Process-based Self-Rewarding Language Models
Shimao Zhang, Xiao Liu, Xin Zhang +4
Large Language Models have demonstrated outstanding performance across various downstream tasks and have been widely applied in multiple scenarios. Human-annotated preference data…