1 paper
Aryan Bhatt, Cody Rushing, Adam Kaufman +5
Control evaluations measure whether monitoring and security protocols for AI systems prevent intentionally subversive AI models from causing harm. Our work presents the first contr…