5 papers
Safety from Honesty in a Disinterested AI Predictor
Yoshua Bengio, Oliver Richardson, Tomáš GavenÄiak +13
As AI systems become more capable, training procedures that optimize for downstream outcomes risk introducing implicit agency: goal-directed behavior that designers never specified…
How does information access affect LLM monitors' ability to detect sabotage?
Rauno Arike, Raja Mehta Moreno, Rohan Subramani +2
Frontier language model agents can exhibit misaligned behaviors, including deception, exploiting reward hacks, and pursuing hidden objectives. To control potentially misaligned age…
Reasoning Under Pressure: How do Training Incentives Influence Chain-of-Thought Monitorability?
Matt MacDermott, Qiyao Wei, Rada Djoneva +1
AI systems that output their reasoning in natural language offer an opportunity for safety -- we can \emph{monitor} their chain of thought (CoT) for undesirable reasoning, such as…
Password-Activated Shutdown Protocols for Misaligned Frontier Agents
Kai Williams, Rohan Subramani, Francis Rhys Ward
Frontier AI developers may fail to align or control highly-capable AI agents. In many cases, it could be useful to have emergency shutdown mechanisms which effectively prevent misa…
Higher-Order Belief in Incomplete Information MAIDs
Jack Foxabbott, Rohan Subramani, Francis Rhys Ward
Multi-agent influence diagrams (MAIDs) are probabilistic graphical models which represent strategic interactions between agents. MAIDs are equivalent to extensive form games (EFGs)…