12 citations · 40 across the 34 of their papers we have counts for
4 papers · 2 filters
What Does It Mean to Break a Distillation Defense?
Lena Libon, Pura Peetathawatchai, Michael Aerni +2
Black-box LLMs (accessible only via API) are vulnerable to distillation attacks, in which an attacker queries the model and trains a student on its outputs. A recent line of work p…
Untrusted Content Masking for Web Agents with Security Guarantees
Kristina Nikolić, Egor Zverev, Javier Rando +3
Defenses that provide security guarantees against prompt injection attacks rely on strict isolation between trusted instructions and untrusted data. In text-based environments such…
Assessing Automated Prompt Injection Attacks in Agentic Environments
David Hofer, Edoardo Debenedetti, Florian Tramèr
Indirect prompt injection poses a critical threat to LLM agents that interact with untrusted external data, yet automated attack methods--proven effective for jailbreaking--remain…
Laundering AI Authority with Adversarial Examples
Jie Zhang, Pura Peetathawatchai, Florian Tramèr +1
Vision-language models (VLMs) are increasingly deployed as trusted authorities -- fact-checking images on social media, comparing products, and moderating content. Users implicitly…