7 papers
Tool Use Enables Undetectable Steganography in Multi-Agent LLM Systems
Jimmy Laurence Rippin, Simon C. Marshall, David Demitri Africa +1
Increasingly autonomous agentic AI systems pose novel multi-agent risks, such as secret collusion via covert communication channels. The natural defence to these collusion attempts…
A Note on the Strategic Confinement Problem
Christian Schroeder de Witt
Lampson's confinement problem asks how to prevent a program that processes confidential information from leaking it to a third party. We introduce the strategic confinement problem…
Detecting Multi-Agent Collusion Through Multi-Agent Interpretability
Aaron Rose, Carissa Cullen, Sahar Abdelnabi +3
As LLM agents are increasingly deployed in multi-agent systems, they introduce risks of covert coordination that may evade standard forms of human oversight. While linear probes on…
A Decision-Theoretic Formalisation of Steganography With Applications to LLM Monitoring
Usman Anwar, Julianna Piskorz, David D. Baek +6
Large language models are beginning to show steganographic capabilities. Such capabilities could allow misaligned models to evade oversight mechanisms. Yet principled methods to de…
Architecture Matters for Multi-Agent Security
Ben Hagag, William L. Anderson, Christian Schroeder de Witt +1
Multi-agent systems (MAS), composed of networks of two or more autonomous AI agents, have become increasingly popular in production deployments, yet introduce security risks that d…
Architecting Resilient LLM Agents: A Guide to Secure Plan-then-Execute Implementations
Ron F. Del Rosario, Klaudia Krawiecka, Christian Schroeder de Witt
As Large Language Model (LLM) agents become increasingly capable of automating complex, multi-step tasks, the need for robust, secure, and predictable architectural patterns is par…