1 paper · 1 filter
Aditya Singh, Gerson Kroiz, Senthooran Rajamanoharan +1
A central goal of safety research is determining whether a model is misaligned. Prior work has largely focused on detecting concerning behavior. But behavior alone does not establi…