14 papers
PiSAs: Benchmarking Contextual Integrity in Multi-User Agentic Systems
Shubham Gupta, Nazanin Mohammadi Sepahvand, Abhinav Kumar +6
As LLM agents evolve from single-user assistants into shared organizational infrastructure, new privacy risks emerge: inappropriate information may not only be exposed through outp…
Janus: a Playground for User-Involved Agentic Permission Management
Natalie Grace Brigham, Eugene Bagdasarian, Tadayoshi Kohno +1
AI agents that autonomously execute tool calls on a user's behalf raise pressing questions about permission management: what role could users play, and what role should they play?…
AI Snitches Get Glitches: Towards Evading Agentic Surveillance
Hyejun Jeong, Dzung Pham, Amir Houmansadr +1
To better assist users with completing challenging tasks, AI agents mediate communications, access data, and interact with different APIs. Many employers (and even nation-states) a…
Targeting World Models to Compromise Robot Learning Pipelines
Ethan Rathbun, Ahmed Agha, Saaduddin Mahmud +3
World models have recently seen a rapid growth in both their popularity and capability as more data efficient tools for generating robot training data or simulating real world envi…
Colosseum: Auditing Collusion in Cooperative Multi-Agent Systems
Mason Nakamura, Abhinav Kumar, Saswat Das +5
Multi-agent systems, where LLM agents communicate through free-form language, enable sophisticated coordination for solving complex cooperative tasks. This surfaces a unique safety…
Understanding Persuasion in Long-Running Agents
Hyejun Jeong, Amir Houmansadr, Shlomo Zilberstein +1
Modern AI agents increasingly combine conversational interaction with autonomous task execution, such as coding and web research, raising a natural question: What happens when an a…