1 paper
Dhananjay Ashok, Ruth-Ann Armstrong, Jonathan May
Language Models (LMs) may acquire harmful knowledge, and yet feign ignorance of these topics when under audit. Inspired by the recent discovery of deception-related behaviour patte…