3 papers
cs.CL2026
MedFabric: Gold Evidence Hides the Difficulty of Word-Level Medical Fabrication Detection
Tung Sum Thomas Kwok, Qian Qian, Xiaofeng Lin +8
Large language models fabricate in medicine, producing fluent statements that are factually wrong, so reliable fabrication detection is a prerequisite for clinical deployment. Repo…
cs.LG2026
Fast and Effective On-policy Distillation from Reasoning Prefixes
Dongxu Zhang, Zhichao Yang, Sepehr Janghorbani +6
On-policy distillation (OPD), which samples trajectories from the student model and supervises them with a teacher at the token level, avoids relying solely on verifiable terminal…
cs.AI2026
Health-SCORE: Towards Scalable Rubrics for Improving Health-LLMs
Zhichao Yang, Sepehr Janghorbani, Dongxu Zhang +6
Rubrics are essential for evaluating open-ended LLM responses, especially in safety-critical domains such as healthcare. However, creating high-quality and domain-specific rubrics…