1 paper · 1 filter
Alexander Gu, Alan Chen
Large language models (LLMs) frequently contradict themselves when the surface form of a logically equivalent question changes. We present a benchmark of 350 question families (1,7…