activity
20242026
most citedDesign Patterns for Securing LLM Agents against Prompt Injections

3 citations · 4 across the 24 of their papers we have counts for

collaborators

24 papers

cs.CR2026

What Does It Mean to Break a Distillation Defense?

Lena Libon, Pura Peetathawatchai, Michael Aerni +2

Black-box LLMs (accessible only via API) are vulnerable to distillation attacks, in which an attacker queries the model and trains a student on its outputs. A recent line of work p…

cs.CR2026

Untrusted Content Masking for Web Agents with Security Guarantees

Kristina Nikolić, Egor Zverev, Javier Rando +3

Defenses that provide security guarantees against prompt injection attacks rely on strict isolation between trusted instructions and untrusted data. In text-based environments such…

cs.CR2026

Assessing Automated Prompt Injection Attacks in Agentic Environments

David Hofer, Edoardo Debenedetti, Florian Tramèr

Indirect prompt injection poses a critical threat to LLM agents that interact with untrusted external data, yet automated attack methods--proven effective for jailbreaking--remain…

cs.LG2026

Reliable and Responsible Foundation Models: A Comprehensive Survey

Xinyu Yang, Junlin Han, Rishi Bommasani +49

Foundation models, including Large Language Models (LLMs), Multimodal Large Language Models (MLLMs), Image Generative Models (i.e, Text-to-Image Models and Image-Editing Models), a…

cs.LG2026

Rethinking Benchmarks for Differentially Private Image Classification

Sabrina Mokhtari, Sara Kodeiri, Shubhankar Mohapatra +2

We revisit benchmarks for differentially private image classification. We suggest a comprehensive set of benchmarks, allowing researchers to evaluate techniques for differentially…

cs.CV2026

Representations of Text and Images Align From Layer One

Evžen Wybitul, Javier Rando, Florian Tramèr +1

We show that for a variety of concepts in adapter-based vision-language models, the representations of their images and their text descriptions are meaningfully aligned from the ve…