3 citations · 4 across the 24 of their papers we have counts for
24 papers
What Does It Mean to Break a Distillation Defense?
Lena Libon, Pura Peetathawatchai, Michael Aerni +2
Black-box LLMs (accessible only via API) are vulnerable to distillation attacks, in which an attacker queries the model and trains a student on its outputs. A recent line of work p…
Untrusted Content Masking for Web Agents with Security Guarantees
Kristina Nikolić, Egor Zverev, Javier Rando +3
Defenses that provide security guarantees against prompt injection attacks rely on strict isolation between trusted instructions and untrusted data. In text-based environments such…
Assessing Automated Prompt Injection Attacks in Agentic Environments
David Hofer, Edoardo Debenedetti, Florian Tramèr
Indirect prompt injection poses a critical threat to LLM agents that interact with untrusted external data, yet automated attack methods--proven effective for jailbreaking--remain…
Reliable and Responsible Foundation Models: A Comprehensive Survey
Xinyu Yang, Junlin Han, Rishi Bommasani +49
Foundation models, including Large Language Models (LLMs), Multimodal Large Language Models (MLLMs), Image Generative Models (i.e, Text-to-Image Models and Image-Editing Models), a…
Rethinking Benchmarks for Differentially Private Image Classification
Sabrina Mokhtari, Sara Kodeiri, Shubhankar Mohapatra +2
We revisit benchmarks for differentially private image classification. We suggest a comprehensive set of benchmarks, allowing researchers to evaluate techniques for differentially…
Representations of Text and Images Align From Layer One
Evžen Wybitul, Javier Rando, Florian Tramèr +1
We show that for a variety of concepts in adapter-based vision-language models, the representations of their images and their text descriptions are meaningfully aligned from the ve…