Showing cs.CLShow all
3 papers · 1 filter
cs.CL2026
RadOT-Eval: Auditable Structured-Evidence Transport for Radiology Report Evaluation
Weixin Liu, Juming Xiong, Yang Li +5
Automatic evaluation is critical for high-stakes text generation, where errors often involve omitted findings, hallucinated content, polarity reversals, location changes, uncertain…
cs.CL2026
Disentangling Prompt Element Level Risk Factors for Hallucinations and Omissions in Mental Health LLM Responses
Congning Ni, Sarvech Qadir, Bryan Steitz +14
Mental health concerns are often expressed outside clinical settings, including in high-distress help seeking, where safety-critical guidance may be needed. Consumer health informa…
cs.CL2026
Coverage-Controlled Preference Mining from Noisy Claim Verification for Evidence-Grounded Generation
Weixin Liu, Congning Ni, Qingyuan Song +4
Evidence-grounded generation produces summaries whose claims should be supported by supplied evidence, but claim-level verifiers provide noisy feedback and can reward models that s…