activity
20242026
most citedPosition: Adversarial ML for LLMs Is Not Making Any Progress

1 citations · 1 across the 2 of their papers we have counts for

collaborators

11 papers

cs.CR2026

Untrusted Content Masking for Web Agents with Security Guarantees

Kristina Nikolić, Egor Zverev, Javier Rando +3

Defenses that provide security guarantees against prompt injection attacks rely on strict isolation between trusted instructions and untrusted data. In text-based environments such…

cs.LG20261 cited

Position: Adversarial ML for LLMs Is Not Making Any Progress

Javier Rando, Jie Zhang, Nicholas Carlini +1

In the past decade, considerable research effort has been devoted to securing machine learning (ML) models that operate in adversarial settings. Yet, progress has been slow even fo…

cs.CR2026

How Vulnerable Are AI Agents to Indirect Prompt Injections? Insights from a Large-Scale Public Competition

Mateusz Dziemian, Maxwell Lin, Xiaohan Fu +28

LLM based agents are increasingly deployed in high stakes settings where they process external data sources such as emails, documents, and code repositories. This creates exposure…

cs.CV2026

Representations of Text and Images Align From Layer One

Evžen Wybitul, Javier Rando, Florian Tramèr +1

We show that for a variety of concepts in adapter-based vision-language models, the representations of their images and their text descriptions are meaningfully aligned from the ve…

cs.LG2025

An Adversarial Perspective on Machine Unlearning for AI Safety

Jakub Łucki, Boyi Wei, Yangsibo Huang +3

Large language models are finetuned to refuse questions about hazardous knowledge, but these protections can often be bypassed. Unlearning methods aim at completely removing hazard…

cs.CR2025

AutoAdvExBench: Benchmarking autonomous exploitation of adversarial example defenses

Nicholas Carlini, Javier Rando, Edoardo Debenedetti +2

We introduce AutoAdvExBench, a benchmark to evaluate if large language models (LLMs) can autonomously exploit defenses to adversarial examples. Unlike existing security benchmarks…