From the 1 of 29 linked papers with an AI index.
29 papers
What Does It Mean to Break a Distillation Defense?
Lena Libon, Pura Peetathawatchai, Michael Aerni +2
The paper examines how to evaluate defenses that add noise to large language model outputs to thwart distillation attacks, proposing a three‑dimensional threat model (query budget,…
Untrusted Content Masking for Web Agents with Security Guarantees
Kristina NikoliÄ, Egor Zverev, Javier Rando +3
Defenses that provide security guarantees against prompt injection attacks rely on strict isolation between trusted instructions and untrusted data. In text-based environments such…
Learning to Inject: Automated Prompt Injection via Reinforcement Learning
Xin Chen, Jie Zhang, Florian Tramèr
Prompt injection is a critical vulnerability in LLM agents, yet the strongest methods still rely on human red-teamers and hand-crafted prompts. Adapting automated jailbreak optimiz…
Assessing Automated Prompt Injection Attacks in Agentic Environments
David Hofer, Edoardo Debenedetti, Florian Tramèr
Indirect prompt injection poses a critical threat to LLM agents that interact with untrusted external data, yet automated attack methods--proven effective for jailbreaking--remain…
CaMeLs Can Use Computers Too: System-level Security for Computer Use Agents
Hanna Foerster, Tom Blanchard, Kristina NikoliÄ +6
AI agents are vulnerable to prompt injection attacks, where malicious content hijacks agent behavior. Among proposed defenses, architectural isolation provides the strongest guaran…
Position: Adversarial ML for LLMs Is Not Making Any Progress
Javier Rando, Jie Zhang, Nicholas Carlini +1
In the past decade, considerable research effort has been devoted to securing machine learning (ML) models that operate in adversarial settings. Yet, progress has been slow even fo…