activity
20162024
most citedDefensive Distillation is Not Robust to Adversarial Examples

238 citations · 411 across the 11 of their papers we have counts for

collaborators

11 papers

cs.CR20241 cited

Privacy Backdoors: Enhancing Membership Inference through Poisoning Pre-trained Models

Yuxin Wen, Leo Marchyok, Sanghyun Hong +3

It is commonplace to produce application-specific models by fine-tuning large pre-trained models using a small bespoke dataset. The widespread availability of foundation model chec…

cs.LG20231 cited

Reverse-Engineering Decoding Strategies Given Blackbox Access to a Language Generation System

Daphne Ippolito, Nicholas Carlini, Katherine Lee +2

Neural language models are increasingly deployed into APIs and websites that allow a user to pass in a prompt and receive generated text. Many of these systems do not reveal genera…

cs.CR202310 cited

Backdoor Attacks for In-Context Learning with Language Models

Nikhil Kandpal, Matthew Jagielski, Florian Tramèr +1

Because state-of-the-art language models are expensive to train, most practitioners must make use of one of the few publicly available language models or language model APIs. This…

cs.CR20233 cited

A LLM Assisted Exploitation of AI-Guardian

Nicholas Carlini

Large language models (LLMs) are now highly capable at a diverse range of tasks. This paper studies whether or not GPT-4, one such LLM, is capable of assisting researchers in the f…

cs.CR20236 cited

Students Parrot Their Teachers: Membership Inference on Model Distillation

Matthew Jagielski, Milad Nasr, Christopher Choquette-Choo +2

Model distillation is frequently proposed as a technique to reduce the privacy leakage of machine learning. These empirical privacy defenses rely on the intuition that distilled ``…

cs.LG20233 cited

Randomness in ML Defenses Helps Persistent Attackers and Hinders Evaluators

Keane Lucas, Matthew Jagielski, Florian Tramèr +2

It is becoming increasingly imperative to design robust ML defenses. However, recent work has found that many defenses that initially resist state-of-the-art attacks can be broken…