4 papers
The Best-Laid SCHEMEs: Coordinated Sabotage and Monitoring in Multi-Agent Systems
Nikolay Radev, Lennart Haas, Benjamin Arnav +1
As agentic coding systems decompose work across multiple model instances, a critical safety question is whether those instances can coordinate to achieve a hidden malicious objecti…
Ensemble Monitoring for AI Control: Diverse Signals Outweigh More Compute
Eugene Koran, Yejun Yun, Samantha Tetef +2
As AI systems are increasingly deployed in autonomous agentic settings at scale, it is important to ensure the actions they take are safe and aligned with user intent. Monitoring a…
Monitoring Monitorability
Melody Y. Guan, Miles Wang, Micah Carroll +9
Observability into the decision making of modern AI systems may be required to safely deploy increasingly capable agents. Monitoring the chain-of-thought (CoT) of today's reasoning…
CoT Red-Handed: Stress Testing Chain-of-Thought Monitoring
Benjamin Arnav, Pablo Bernabeu-Pérez, Nathan Helm-Burger +3
As AI models are deployed with increasing autonomy, it is important to ensure they do not take harmful actions unnoticed. As a potential mitigation, we investigate Chain-of-Thought…