1 citations · 1 across the 2 of their papers we have counts for
11 papers
Untrusted Content Masking for Web Agents with Security Guarantees
Kristina NikoliÄ, Egor Zverev, Javier Rando +3
Defenses that provide security guarantees against prompt injection attacks rely on strict isolation between trusted instructions and untrusted data. In text-based environments such…
Position: Adversarial ML for LLMs Is Not Making Any Progress
Javier Rando, Jie Zhang, Nicholas Carlini +1
In the past decade, considerable research effort has been devoted to securing machine learning (ML) models that operate in adversarial settings. Yet, progress has been slow even fo…
How Vulnerable Are AI Agents to Indirect Prompt Injections? Insights from a Large-Scale Public Competition
Mateusz Dziemian, Maxwell Lin, Xiaohan Fu +28
LLM based agents are increasingly deployed in high stakes settings where they process external data sources such as emails, documents, and code repositories. This creates exposure…
Representations of Text and Images Align From Layer One
Evžen Wybitul, Javier Rando, Florian Tramèr +1
We show that for a variety of concepts in adapter-based vision-language models, the representations of their images and their text descriptions are meaningfully aligned from the ve…
An Adversarial Perspective on Machine Unlearning for AI Safety
Jakub Åucki, Boyi Wei, Yangsibo Huang +3
Large language models are finetuned to refuse questions about hazardous knowledge, but these protections can often be bypassed. Unlearning methods aim at completely removing hazard…
AutoAdvExBench: Benchmarking autonomous exploitation of adversarial example defenses
Nicholas Carlini, Javier Rando, Edoardo Debenedetti +2
We introduce AutoAdvExBench, a benchmark to evaluate if large language models (LLMs) can autonomously exploit defenses to adversarial examples. Unlike existing security benchmarks…