3 papers
cs.CL2026
This Treatment Works, Right? Evaluating LLM Sensitivity to Patient Question Framing in Medical QA
Hye Sun Yun, Geetika Kapoor, Michael Mackert +4
Patients are increasingly turning to large language models (LLMs) with medical questions that are complex and difficult to articulate clearly. However, LLMs are sensitive to prompt…
cs.CL2026
Faithfulness vs. Safety: Evaluating LLM Behavior Under Counterfactual Medical Evidence
Kaijie Mo, Siddhartha Venkatayogi, Chantal Shaib +4
In high-stakes domains like medicine, it may be generally desirable for models to faithfully adhere to the context provided. But what happens if the context does not align with mod…
cs.CL2025
Decide less, communicate more: On the construct validity of end-to-end fact-checking in medicine
Sebastian Joseph, Lily Chen, Barry Wei +6
Technological progress has led to concrete advancements in tasks that were regarded as challenging, such as automatic fact-checking. Interest in adopting these systems for public h…