Showing cs.CLShow all
2 papers · 1 filter
cs.CL2026
FramingQA: Does the Question Shape the Answer? Measuring the Compositional Framing Effect
Hazel H. Kim, Andrew M. Bean, Guilherme Affonso Ferreira de Camargo +10
We introduce FramingQA, a benchmark that measures the model sensitivity to question framing across law, medicine, finance, and robotic simulations. Large language models (LLMs) oft…
cs.CL2026
Measuring Epistemic Resilience of LLMs Under Misleading Medical Context
Hongjian Zhou, Xinyu Zou, Jinge Wu +19
Large language models (LLMs) now reach expert-level scores on medical licensing exams, encouraging the assumption that high scores imply safe medical judgment while patients increa…