18 papers
Reasoning Error from Known Fact: Step-Level Self-Consistency Group Relative Policy Optimization for LLM
Xiaomeng Hu, Jiaqi Hu, Hao Chen +4
With the rapid advancement of large language models (LLMs), modern systems not only possess strong foundational capabilities and extensive knowledge, but can also solve complex pro…
Semantically Similar, Logically Distinct: Diagnosing the Semantic-Answerability Gap in Table RAG
Jiaming Tian, Liyao Li, Wentao Ye +5
Tables are a critical knowledge source in retrieval-augmented generation (RAG), but a retrieved table may lack sufficient evidence to answer a query, a property we call answerabili…
Purified OPSD: On-Policy Self-Distillation Without Losing How to Think
Zhanming Shen, Jintao Tong, Shaotian Yan +9
On-policy self-distillation (OPSD) has emerged as a promising paradigm for improving LLM reasoning, where a privileged teacher with access to reference solutions provides token-lev…
Eliciting Complex Spatial Reasoning in MLLMs through Wide-Baseline Matching
Hao Zhong, Muzhi Zhu, Shenyan Zeng +8
Wide-baseline matching (WBM) requires integrating geometric understanding, viewpoint changes, fine-grained perception, and occlusion reasoning, making it a challenging testbed for…
FLaG: Fine-Grained Latent Grouping for Hallucination Detection
Wentao Ye, Liyao Li, Zhiqing Xiao +6
Hallucinations in large language models (LLMs) arise from heterogeneous failure mechanisms, making reliable detection difficult for any single global uncertainty score. In this wor…
Training-Trajectory-Aware Token Selection
Zhanming Shen, Jiaqi Hu, Zeyu Qin +7
Efficient distillation is a key pathway for converting expensive reasoning capability into deployable efficiency, yet in the frontier regime where the student already has strong re…