1 paper
Maheep Chaudhary, Ian Su, Nikhil Hooda +6
Large language models (LLMs) can internally distinguish between evaluation and deployment contexts, a behaviour known as \emph{evaluation awareness}. This undermines AI safety eval…