1 paper · 1 filter
Ziyi Ding, Xiao-Ping Zhang
We find a mismatch between what large language models encode about a causal question and what they answer. On anti-commonsense CLadder items, a fixed linear probe recovers the evid…