15 citations · 28 across the 3 of their papers we have counts for
3 papers
Rethinking Backdoor Attacks
Alaa Khaddaj, Guillaume Leclerc, Aleksandar Makelov +4
In a backdoor attack, an adversary inserts maliciously constructed backdoor examples into a training set to make the resulting model vulnerable to manipulation. Defending against s…
TRAK: Attributing Model Behavior at Scale
Sung Min Park, Kristian Georgiev, Andrew Ilyas +2
The goal of data attribution is to trace model predictions back to training data. Despite a long line of work towards this goal, existing approaches to data attribution tend to for…
Raising the Cost of Malicious AI-Powered Image Editing
Hadi Salman, Alaa Khaddaj, Guillaume Leclerc +2
We present an approach to mitigating the risks of malicious image editing posed by large diffusion models. The key idea is to immunize images so as to make them resistant to manipu…