3 papers
cs.LG2024
Subversion Strategy Eval: Can language models statelessly strategize to subvert control protocols?
Alex Mallen, Charlie Griffin, Misha Wagner +2
An AI control protocol is a plan for usefully deploying AI systems that aims to prevent an AI from intentionally causing some unacceptable outcome. This paper investigates how well…
cs.AI2024
Games for AI Control: Models of Safety Evaluations of AI Deployment Protocols
Charlie Griffin, Louis Thomson, Buck Shlegeris +1
To evaluate the safety and usefulness of deployment protocols for untrusted AIs, AI Control uses a red-teaming exercise played between a protocol designer and an adversary. This pa…
cs.LG2023
Reinforcement Learning Fine-tuning of Language Models is Biased Towards More Extractable Features
Diogo Cruz, Edoardo Pona, Alex Holness-Tofts +4
Many capable large language models (LLMs) are developed via self-supervised pre-training followed by a reinforcement-learning fine-tuning phase, often based on human or AI feedback…