activity
20242026
collaborators

8 papers

cs.CL2026

Coverage-Controlled Preference Mining from Noisy Claim Verification for Evidence-Grounded Generation

Weixin Liu, Congning Ni, Qingyuan Song +4

Evidence-grounded generation produces summaries whose claims should be supported by supplied evidence, but claim-level verifiers provide noisy feedback and can reward models that s…

cs.CL2026

RadOT-Eval: Auditable Structured-Evidence Transport for Radiology Report Evaluation

Weixin Liu, Juming Xiong, Yang Li +5

Automatic evaluation is critical for high-stakes text generation, where errors often involve omitted findings, hallucinated content, polarity reversals, location changes, uncertain…

cs.CL2026

MHGraphBench: Knowledge Graph-Grounded Benchmarking of Mental Health Knowledge in Large Language Models

Weixin Liu, Congning Ni, Shelagh A. Mulvaney +4

Large language models (LLMs) are increasingly used in the mental health domain, yet it remains unclear how well they capture related biomedical knowledge and how reliably they appl…

cs.CL2026

Blending Human and LLM Expertise to Detect Hallucinations and Omissions in Mental Health Chatbot Responses

Khizar Hussain, Bradley A. Malin, Zhijun Yin +2

As LLM-powered chatbots are increasingly deployed in mental health services, detecting hallucinations and omissions has become critical for user safety. However, state-of-the-art L…

cs.CL2026

Disentangling Prompt Element Level Risk Factors for Hallucinations and Omissions in Mental Health LLM Responses

Congning Ni, Sarvech Qadir, Bryan Steitz +14

Mental health concerns are often expressed outside clinical settings, including in high-distress help seeking, where safety-critical guidance may be needed. Consumer health informa…

cs.CL2025

Judging with Confidence: Calibrating Autoraters to Preference Distributions

Zhuohang Li, Xiaowei Li, Chengyu Huang +11

The alignment of large language models (LLMs) with human values increasingly relies on using other LLMs as automated judges, or ``autoraters''. However, their reliability is limite…