2 papers
cs.CL2026
Judging by the Cover: Cleaning LLM Truthfulness Benchmarks to Avoid Surface-Level Feature Leakage
Foad Namjoo, Remy Ogasawara, Amirali Abdullah +3
Binary-choice truth benchmarks ask models to choose between a correct and an incorrect answer, but if the two answers differ systematically in surface-level features, models can ex…
cs.LG2026
Understanding and Mitigating Dataset Corruption in LLM Steering
Cullen Anderson, Narmeen Oozeer, Foad Namjoo +3
Contrastive steering has been shown as a simple and effective method to adjust the generative behavior of LLMs at inference time. It uses examples of prompt responses with and with…