5 papers
Is your algorithm unlearning or untraining?
Eleni Triantafillou, Ahmed Imtiaz Humayun, Monica Ribero +3
As models are getting larger and are trained on increasing amounts of data, there has been an explosion of interest into how we can ``delete'' specific data points or behaviours fr…
Detoxifying LLMs via Representation Erasure-Based Preference Optimization
Nazanin Mohammadi Sepahvand, Eleni Triantafillou, Hugo Larochelle +3
Large language models (LLMs) trained on webscale data can produce toxic outputs, raising concerns for safe deployment. Prior defenses, based on applications of DPO, NPO, and simila…
Step-resolved data attribution for looped transformers
Georgios Kaissis, David Mildenberger, Juan Felipe Gomez +2
We study how individual training examples shape the internal computation of looped transformers, where a shared block is applied for recurrent iterations to enable latent reas…
Position: Capability Control Should be a Separate Goal From Alignment
Shoaib Ahmed Siddiqui, Eleni Triantafillou, David Krueger +1
Foundation models are trained on broad data distributions, yielding generalist capabilities that enable many downstream applications but also expand the space of potential misuse a…
Your Privacy Depends on Others: Collusion Vulnerabilities in Individual Differential Privacy
Johannes Kaiser, Alexander Ziller, Eleni Triantafillou +2
Individual Differential Privacy (iDP) promises users control over their privacy, but this promise can be broken in practice. We reveal a previously overlooked vulnerability in samp…