From the 1 of 1 linked paper with an AI index.
1 paper
Alexander Gu, Alan Chen
The paper introduces CRTBench, a benchmark of 350 question families to test whether large language models give consistent answers across controlled reformulations such as contrapos…