2 papers
cs.AI2026
When can we trust untrusted monitoring? A safety case sketch across collusion strategies
Nelson Gardner-Challis, Jonathan Bostock, Georgiy Kozhevnikov +4
AIs are increasingly being deployed with greater autonomy and capabilities, which increases the risk that a misaligned AI may be able to cause catastrophic harm. Untrusted monitori…
cs.LG2024
Subversion Strategy Eval: Can language models statelessly strategize to subvert control protocols?
Alex Mallen, Charlie Griffin, Misha Wagner +2
An AI control protocol is a plan for usefully deploying AI systems that aims to prevent an AI from intentionally causing some unacceptable outcome. This paper investigates how well…