7 papers
Language Generation with Replay: A Learning-Theoretic View of Model Collapse
Giorgio Racca, Michal Valko, Amartya Sanyal
As scaling laws push the training of frontier large language models (LLMs) toward ever-growing data requirements, training pipelines are approaching a regime where much of the publ…
Less Noise, Same Certificate: Retain Sensitivity for Unlearning
Carolin Heinzler, Kasra Malihi, Amartya Sanyal
Certified machine unlearning aims to provably remove the influence of a deletion set from a model trained on a dataset , by producing an unlearned output that is statistical…
Fairness for the People, by the People: Minority Collective Action
Omri Ben-Dov, Samira Samadi, Amartya Sanyal +1
Machine learning models often preserve biases present in training data, leading to unfair treatment of certain minority groups. Despite an array of existing firm-side bias mitigati…
Delta-Influence: Unlearning Poisons via Influence Functions
Wenjie Li, Jiawei Li, Pengcheng Zeng +3
Addressing data integrity challenges, such as unlearning the effects of data poisoning after model training, is necessary for the reliable deployment of machine learning models. St…
An Iterative Algorithm for Differentially Private -PCA with Adaptive Noise
Johanna Düngler, Amartya Sanyal
Given i.i.d. random matrices that share a common expectation , the objective of Differentially Private Stochastic PCA is to identify a sub…
Provable unlearning in topic modeling and downstream tasks
Stanley Wei, Sadhika Malladi, Sanjeev Arora +1
Machine unlearning algorithms are increasingly important as legal concerns arise around the provenance of training data, but verifying the success of unlearning is often difficult.…