14 papers
Firewalls to Secure Dynamic LLM Agentic Networks
Sahar Abdelnabi, Amr Gomaa, Eugene Bagdasarian +2
The emergence of agent-to-agent communication protocols mirrors the early internet: powerful connectivity with minimal security infrastructure. When AI agents communicate on behalf…
Models That Know How Evaluations Are Designed Score Safer
Katharina Deckenbach, Haritz Puerto, Jonas Geiping +1
The validity of AI safety evaluations depends on models behaving consistently across controlled and deployment settings. Prior work has identified test-time contextual cues, such a…
Decomposing and Measuring Evaluation Awareness
Changling Li, Terry Jingchen Zhang, Jie Zhang +3
Frontier language models sometimes recognize that they are being evaluated and adjust their behavior, undermining validity of benchmark results. Yet the field studies it without a…
Colosseum: Auditing Collusion in Cooperative Multi-Agent Systems
Mason Nakamura, Abhinav Kumar, Saswat Das +5
Multi-agent systems, where LLM agents communicate through free-form language, enable sophisticated coordination for solving complex cooperative tasks. This surfaces a unique safety…
Measuring Security Without Fooling Ourselves: Why Benchmarking Agents Is Hard
Sahar Abdelnabi, Chris Hicks, Konrad Rieck +1
The benchmarks used to evaluate AI agents in security-critical roles suffer from crucial weaknesses. Building on recent empirical evidence, we characterize three core challenges th…
Hidden in Memory: Sleeper Memory Poisoning in LLM Agents
Sidharth Pulipaka, Stanislau Hlebik, Leonidas Raghav +4
Large language models are increasingly augmented with persistent memory, allowing assistants to store user-specific information across sessions for personalization and continuity.…