3 papers
cs.CL2026
MedFabric: Gold Evidence Hides the Difficulty of Word-Level Medical Fabrication Detection
Tung Sum Thomas Kwok, Qian Qian, Xiaofeng Lin +8
Large language models fabricate in medicine, producing fluent statements that are factually wrong, so reliable fabrication detection is a prerequisite for clinical deployment. Repo…
cs.CL2025
Play Favorites: A Statistical Method to Measure Self-Bias in LLM-as-a-Judge
Evangelia Spiliopoulou, Riccardo Fogliato, Hanna Burnsky +4
Large language models (LLMs) can serve as judges that offer rapid and reliable assessments of other LLM outputs. However, models may systematically assign overly favorable ratings…
cs.CL2024
Leveraging LLMs for Dialogue Quality Measurement
Jinghan Jia, Abi Komma, Timothy Leffel +5
In task-oriented conversational AI evaluation, unsupervised methods poorly correlate with human judgments, and supervised approaches lack generalization. Recent advances in large l…