6 citations · 6 across the 3 of their papers we have counts for
1 paper · 1 filter
Xian Sun, Wei Chow, Yingshuo Wang +4
Language models increasingly condition their answers on external signals, and a single misleading one can turn a correct answer wrong. The obvious remedy, training models to resist…