3 papers
cs.CL2026
Faithfulness vs. Safety: Evaluating LLM Behavior Under Counterfactual Medical Evidence
Kaijie Mo, Siddhartha Venkatayogi, Chantal Shaib +4
In high-stakes domains like medicine, it may be generally desirable for models to faithfully adhere to the context provided. But what happens if the context does not align with mod…
cs.CL2026
FAITH: Factuality Alignment through Integrating Trustworthiness and Honestness
Xiaoning Dong, Chengyan Wu, Yajie Wen +5
Large Language Models (LLMs) can generate factually inaccurate content even if they have corresponding knowledge, which critically undermines their reliability. Existing approaches…
cs.CL2026
This Treatment Works, Right? Evaluating LLM Sensitivity to Patient Question Framing in Medical QA
Hye Sun Yun, Geetika Kapoor, Michael Mackert +4
Patients are increasingly turning to large language models (LLMs) with medical questions that are complex and difficult to articulate clearly. However, LLMs are sensitive to prompt…