1 paper · 1 filter
Foad Namjoo, Remy Ogasawara, Amirali Abdullah +3
Binary-choice truth benchmarks ask models to choose between a correct and an incorrect answer, but if the two answers differ systematically in surface-level features, models can ex…