12 papers
What to Forget in Unlearning? Forget Set Curation for Language Models
Animesh Jha, Arpandeep Khatua, Youssef Allouah +1
Machine unlearning aims to remove targeted data or behaviors from a trained model without retraining from scratch. Yet most evaluations assume that the examples to forget are alrea…
The Utility and Complexity of in- and out-of-Distribution Machine Unlearning
Youssef Allouah, Joshua Kazdan, Rachid Guerraoui +1
Machine unlearning, the process of selectively removing data from trained models, is increasingly crucial for addressing privacy concerns and knowledge gaps post-deployment. Despit…
The Distillation Game: Adaptive Attacks & Efficient Defenses
Youssef Allouah, Mahdi Haghifam, Sanmi Koyejo +1
Distillation attacks create a deployment trade-off for model providers: the same outputs that make a model more useful can also make it easier to imitate. We study this trade-off t…
Byzantine-Robust and Communication-Efficient Distributed Training: Compressive and Cyclic Gradient Coding
Chengxi Li, Youssef Allouah, Rachid Guerraoui +2
In this paper, we study the problem of distributed training (DT) under Byzantine attacks with communication constraints. While prior work has developed various robust aggregation r…
Distributional Machine Unlearning via Selective Data Removal
Youssef Allouah, Rachid Guerraoui, Sanmi Koyejo
Machine learning systems increasingly face requirements to remove entire domains of information--such as toxic language or biases--rather than individual user data. This task prese…
Efficient Prediction of Pass@k Scaling in Large Language Models
Joshua Kazdan, Rylan Schaeffer, Youssef Allouah +4
Assessing the capabilities and risks of frontier AI systems is a critical area of research, and recent work has shown that repeated sampling from models can dramatically increase b…