1 paper · 1 filter
Igor Ivanov, David Demitri Africa
Large language models can recognize when they are being evaluated (evaluation awareness) and behave differently because of that, which undermines the validity of safety and alignme…