4 papers
How does information access affect LLM monitors' ability to detect sabotage?
Rauno Arike, Raja Mehta Moreno, Rohan Subramani +2
Frontier language model agents can exhibit misaligned behaviors, including deception, exploiting reward hacks, and pursuing hidden objectives. To control potentially misaligned age…
Password-Activated Shutdown Protocols for Misaligned Frontier Agents
Kai Williams, Rohan Subramani, Francis Rhys Ward
Frontier AI developers may fail to align or control highly-capable AI agents. In many cases, it could be useful to have emergency shutdown mechanisms which effectively prevent misa…
Higher-Order Belief in Incomplete Information MAIDs
Jack Foxabbott, Rohan Subramani, Francis Rhys Ward
Multi-agent influence diagrams (MAIDs) are probabilistic graphical models which represent strategic interactions between agents. MAIDs are equivalent to extensive form games (EFGs)…
The Partially Observable Off-Switch Game
Andrew Garber, Rohan Subramani, Linus Luu +3
A wide variety of goals could cause an AI to disable its off switch because "you can't fetch the coffee if you're dead" (Russell 2019). Prior theoretical work on this shutdown prob…