1 paper
Olli Järviniemi, Oliver Makins, Jacob Merizian +2
Motivated by loss of control risks from misaligned AI systems, we develop and apply methods for measuring language models' propensity for unsanctioned behaviour. We contribute thre…