works on

From the 1 of 29 linked papers with an AI index.

activity
20242026
collaborators

29 papers

cs.CR2026

What Does It Mean to Break a Distillation Defense?

Lena Libon, Pura Peetathawatchai, Michael Aerni +2

The paper examines how to evaluate defenses that add noise to large language model outputs to thwart distillation attacks, proposing a three‑dimensional threat model (query budget,…

cs.CR2026

Untrusted Content Masking for Web Agents with Security Guarantees

Kristina Nikolić, Egor Zverev, Javier Rando +3

Defenses that provide security guarantees against prompt injection attacks rely on strict isolation between trusted instructions and untrusted data. In text-based environments such…

cs.LG2026

Learning to Inject: Automated Prompt Injection via Reinforcement Learning

Xin Chen, Jie Zhang, Florian Tramèr

Prompt injection is a critical vulnerability in LLM agents, yet the strongest methods still rely on human red-teamers and hand-crafted prompts. Adapting automated jailbreak optimiz…

cs.CR2026

Assessing Automated Prompt Injection Attacks in Agentic Environments

David Hofer, Edoardo Debenedetti, Florian Tramèr

Indirect prompt injection poses a critical threat to LLM agents that interact with untrusted external data, yet automated attack methods--proven effective for jailbreaking--remain…

cs.AI2026

CaMeLs Can Use Computers Too: System-level Security for Computer Use Agents

Hanna Foerster, Tom Blanchard, Kristina Nikolić +6

AI agents are vulnerable to prompt injection attacks, where malicious content hijacks agent behavior. Among proposed defenses, architectural isolation provides the strongest guaran…

cs.LG2026

Position: Adversarial ML for LLMs Is Not Making Any Progress

Javier Rando, Jie Zhang, Nicholas Carlini +1

In the past decade, considerable research effort has been devoted to securing machine learning (ML) models that operate in adversarial settings. Yet, progress has been slow even fo…