collaborators

12 papers

cs.CL2026

Reasoning Error from Known Fact: Step-Level Self-Consistency Group Relative Policy Optimization for LLM

Xiaomeng Hu, Jiaqi Hu, Hao Chen +4

With the rapid advancement of large language models (LLMs), modern systems not only possess strong foundational capabilities and extensive knowledge, but can also solve complex pro…

cs.AI2026

Purified OPSD: On-Policy Self-Distillation Without Losing How to Think

Zhanming Shen, Jintao Tong, Shaotian Yan +9

On-policy self-distillation (OPSD) has emerged as a promising paradigm for improving LLM reasoning, where a privileged teacher with access to reference solutions provides token-lev…

cs.AI2026

Momentum for Reasoning: Dense Intrinsic Signals in Policy Optimization

Hao Chen, Zhanming Shen, Liyao Li +8

Reinforcement learning with verifiable rewards (RLVR) has emerged as a powerful paradigm for eliciting long-chain reasoning in large language models. However, existing methods base…

cs.LG2026

FLaG: Fine-Grained Latent Grouping for Hallucination Detection

Wentao Ye, Liyao Li, Zhiqing Xiao +6

Hallucinations in large language models (LLMs) arise from heterogeneous failure mechanisms, making reliable detection difficult for any single global uncertainty score. In this wor…

cs.CL2026

Training-Trajectory-Aware Token Selection

Zhanming Shen, Jiaqi Hu, Zeyu Qin +7

Efficient distillation is a key pathway for converting expensive reasoning capability into deployable efficiency, yet in the frontier regime where the student already has strong re…

cs.LG2026

From Parameters to Data: A Task-Parameter-Guided Fine-Tuning Pipeline for Efficient LLM Alignment

Hao Chen, Qi Zhang, Liyao Li +7

Adapting Large Language Models (LLMs) to specialized domains typically incurs high data and computational overhead. While prior efficiency efforts have largely treated data selecti…