1 citations · 1 across the 1 of their papers we have counts for
3 papers
Games for AI Control: Models of Safety Evaluations of AI Deployment Protocols
Charlie Griffin, Louis Thomson, Buck Shlegeris +1
To evaluate the safety and usefulness of deployment protocols for untrusted AIs, AI Control uses a red-teaming exercise played between a protocol designer and an adversary. This pa…
When can we trust untrusted monitoring? A safety case sketch across collusion strategies
Nelson Gardner-Challis, Jonathan Bostock, Georgiy Kozhevnikov +4
AIs are increasingly being deployed with greater autonomy and capabilities, which increases the risk that a misaligned AI may be able to cause catastrophic harm. Untrusted monitori…
Subversion Strategy Eval: Can language models statelessly strategize to subvert control protocols?
Alex Mallen, Charlie Griffin, Misha Wagner +2
An AI control protocol is a plan for usefully deploying AI systems that aims to prevent an AI from intentionally causing some unacceptable outcome. This paper investigates how well…