collaborators

5 papers

cs.AI2026

Safety from Honesty in a Disinterested AI Predictor

Yoshua Bengio, Oliver Richardson, Tomáš Gavenčiak +13

As AI systems become more capable, training procedures that optimize for downstream outcomes risk introducing implicit agency: goal-directed behavior that designers never specified…

cs.AI2026

How does information access affect LLM monitors' ability to detect sabotage?

Rauno Arike, Raja Mehta Moreno, Rohan Subramani +2

Frontier language model agents can exhibit misaligned behaviors, including deception, exploiting reward hacks, and pursuing hidden objectives. To control potentially misaligned age…

cs.AI2025

Reasoning Under Pressure: How do Training Incentives Influence Chain-of-Thought Monitorability?

Matt MacDermott, Qiyao Wei, Rada Djoneva +1

AI systems that output their reasoning in natural language offer an opportunity for safety -- we can \emph{monitor} their chain of thought (CoT) for undesirable reasoning, such as…

cs.CR2025

Password-Activated Shutdown Protocols for Misaligned Frontier Agents

Kai Williams, Rohan Subramani, Francis Rhys Ward

Frontier AI developers may fail to align or control highly-capable AI agents. In many cases, it could be useful to have emergency shutdown mechanisms which effectively prevent misa…

cs.AI2025

Higher-Order Belief in Incomplete Information MAIDs

Jack Foxabbott, Rohan Subramani, Francis Rhys Ward

Multi-agent influence diagrams (MAIDs) are probabilistic graphical models which represent strategic interactions between agents. MAIDs are equivalent to extensive form games (EFGs)…