16 papers
Reasoning Error from Known Fact: Step-Level Self-Consistency Group Relative Policy Optimization for LLM
Xiaomeng Hu, Jiaqi Hu, Hao Chen +4
With the rapid advancement of large language models (LLMs), modern systems not only possess strong foundational capabilities and extensive knowledge, but can also solve complex pro…
Purified OPSD: On-Policy Self-Distillation Without Losing How to Think
Zhanming Shen, Jintao Tong, Shaotian Yan +9
On-policy self-distillation (OPSD) has emerged as a promising paradigm for improving LLM reasoning, where a privileged teacher with access to reference solutions provides token-lev…
Verifiable Geometry Problem Solving: Solver-Driven Autoformalization and Theorem Proposing
Can Li, Ting Zhang, Junbo Zhao +1
Geometry Problem Solving have increasingly adopt the neuro-symbolic paradigm, combining neural intuition with symbolic rigor. However, current frameworks suffer from severe bottlen…
OPRD: On-Policy Representation Distillation
Shenzhi Yang, Guangcheng Zhu, Bowen Song +8
On-policy distillation (OPD) supervises the student exclusively in the output space by matching next-token distributions. This paradigm suffers from two limitations: (i) a high-var…
Non-Parametric Structural Priors for Geometry Theorem Prediction
Junbo Zhao, Ting Zhang, Can Li +3
Multi-step theorem prediction is a central challenge in geometry problem solving. Existing neural-symbolic approaches rely heavily on supervised parametric models, which exhibit li…
Training-Trajectory-Aware Token Selection
Zhanming Shen, Jiaqi Hu, Zeyu Qin +7
Efficient distillation is a key pathway for converting expensive reasoning capability into deployable efficiency, yet in the frontier regime where the student already has strong re…