collaborators

16 papers

cs.CL2026

Reasoning Error from Known Fact: Step-Level Self-Consistency Group Relative Policy Optimization for LLM

Xiaomeng Hu, Jiaqi Hu, Hao Chen +4

With the rapid advancement of large language models (LLMs), modern systems not only possess strong foundational capabilities and extensive knowledge, but can also solve complex pro…

cs.AI2026

Purified OPSD: On-Policy Self-Distillation Without Losing How to Think

Zhanming Shen, Jintao Tong, Shaotian Yan +9

On-policy self-distillation (OPSD) has emerged as a promising paradigm for improving LLM reasoning, where a privileged teacher with access to reference solutions provides token-lev…

cs.AI2026

Verifiable Geometry Problem Solving: Solver-Driven Autoformalization and Theorem Proposing

Can Li, Ting Zhang, Junbo Zhao +1

Geometry Problem Solving have increasingly adopt the neuro-symbolic paradigm, combining neural intuition with symbolic rigor. However, current frameworks suffer from severe bottlen…

cs.LG2026

OPRD: On-Policy Representation Distillation

Shenzhi Yang, Guangcheng Zhu, Bowen Song +8

On-policy distillation (OPD) supervises the student exclusively in the output space by matching next-token distributions. This paradigm suffers from two limitations: (i) a high-var…

cs.AI2026

Non-Parametric Structural Priors for Geometry Theorem Prediction

Junbo Zhao, Ting Zhang, Can Li +3

Multi-step theorem prediction is a central challenge in geometry problem solving. Existing neural-symbolic approaches rely heavily on supervised parametric models, which exhibit li…

cs.CL2026

Training-Trajectory-Aware Token Selection

Zhanming Shen, Jiaqi Hu, Zeyu Qin +7

Efficient distillation is a key pathway for converting expensive reasoning capability into deployable efficiency, yet in the frontier regime where the student already has strong re…