David D. Baek, Xinnuo Li, Anay Gupta +4
The increasing situational awareness of language models raises safety concerns: models might be aware when they are evaluated, and adjust their behavior to evade monitoring and res…