1 citations · 2 across the 17 of their papers we have counts for
Showing 2026 · cs.CLShow all
2 papers · 2 filters
cs.CL2026
ROSD: Reflective On-Policy Self-Distillation for Language Model Reasoning across Domains
Ziqi Zhao, Xinyu Ma, Liu Yang +6
On-policy self-distillation (OPSD) improves the reasoning performance of large language models (LLMs) by providing dense token-level supervision for on-policy rollouts. However, ex…
cs.CL2026
Micro-Macro Retrieval: Reducing Long-Form Hallucination in Large Language Models
Yujie Feng, Jian Li, Zhihan Zhou +7
Large Language Models (LLMs) achieve impressive performance across many tasks but remain prone to hallucination, especially in long-form generation where redundant retrieved contex…