1 paper match
Alexander Gu, Alan Chen
The paper introduces CRTBench, a benchmark of 350 question families to test whether large language models give consistent answers across controlled reformulations such as contrapos…