4 papers · 1 filter
FramingQA: Does the Question Shape the Answer? Measuring the Compositional Framing Effect
Hazel H. Kim, Andrew M. Bean, Guilherme Affonso Ferreira de Camargo +10
We introduce FramingQA, a benchmark that measures the model sensitivity to question framing across law, medicine, finance, and robotic simulations. Large language models (LLMs) oft…
LINGOLY-TOO: Disentangling Reasoning from Knowledge with Templatised Orthographic Obfuscation
Jude Khouja, Lingyi Yang, Karolina Korgul +6
Frontier language models demonstrate increasing ability at solving reasoning problems, but their performance is often inflated by circumventing reasoning and instead relying on the…
Evaluating Fine-Tuning Efficiency of Human-Inspired Learning Strategies in Medical Question Answering
Yushi Yang, Andrew M. Bean, Robert McCraith +1
Fine-tuning Large Language Models (LLMs) incurs considerable training costs, driving the need for data-efficient training with optimised data ordering. Human-inspired strategies of…
Do Large Language Models have Shared Weaknesses in Medical Question Answering?
Andrew M. Bean, Karolina Korgul, Felix Krones +2
Large language models (LLMs) have made rapid improvement on medical benchmarks, but their unreliability remains a persistent challenge for safe real-world uses. To design for the u…