733 citations · 1.9k across the 46 of their papers we have counts for
3 papers · 2 filters
Scalable Extraction of Training Data from (Production) Language Models
Milad Nasr, Nicholas Carlini, Jonathan Hayase +7
This paper studies extractable memorization: training data that an adversary can efficiently extract by querying a machine learning model without prior knowledge of the training da…
Evaluating Superhuman Models with Consistency Checks
Lukas Fluri, Daniel Paleka, Florian Tramèr
If machine learning models were to achieve superhuman abilities at various reasoning or decision-making tasks, how would we go about evaluating such models, given that humans would…
Randomness in ML Defenses Helps Persistent Attackers and Hinders Evaluators
Keane Lucas, Matthew Jagielski, Florian Tramèr +2
It is becoming increasingly imperative to design robust ML defenses. However, recent work has found that many defenses that initially resist state-of-the-art attacks can be broken…