works on

From the 1 of 8 linked papers with an AI index.

activity
20242026
collaborators
Showing cs.CRShow all

5 papers · 1 filter

cs.CR2026

What Does It Mean to Break a Distillation Defense?

Lena Libon, Pura Peetathawatchai, Michael Aerni +2

The paper examines how to evaluate defenses that add noise to large language model outputs to thwart distillation attacks, proposing a three‑dimensional threat model (query budget,…

cs.CR2026

Large-scale online deanonymization with LLMs

Simon Lermen, Daniel Paleka, Joshua Swanson +3

We show that large language models can be used to perform at-scale deanonymization. With full Internet access, our agent can re-identify Hacker News users and Anthropic Interviewer…

cs.CR2024

Stealing Part of a Production Language Model

Nicholas Carlini, Daniel Paleka, Krishnamurthy Dj Dvijotham +11

We introduce the first model-stealing attack that extracts precise, nontrivial information from black-box production language models like OpenAI's ChatGPT or Google's PaLM-2. Speci…

cs.CR2024

Dataset and Lessons Learned from the 2024 SaTML LLM Capture-the-Flag Competition

Edoardo Debenedetti, Javier Rando, Daniel Paleka +18

Large language model systems face important security risks from maliciously crafted messages that aim to overwrite the system's original instructions or leak private data. To study…

cs.CR2024

Poisoning Web-Scale Training Datasets is Practical

Nicholas Carlini, Matthew Jagielski, Christopher A. Choquette-Choo +6

Deep learning models are often trained on distributed, web-scale datasets crawled from the internet. In this paper, we introduce two new dataset poisoning attacks that intentionall…