1 paper · 1 filter
Hazel H. Kim, Andrew M. Bean, Guilherme Affonso Ferreira de Camargo +10
We introduce FramingQA, a benchmark that measures the model sensitivity to question framing across law, medicine, finance, and robotic simulations. Large language models (LLMs) oft…