9 papers
De-attribute to Forget for LLM Unlearning
Xinyang Lu, Jiabao Pan, Rachael Hwee Ling Sim +3
The rapid development of large language models (LLMs) has raised concerns on the use of inappropriate data for training, which has led to a growing interest in LLM unlearning. Many…
How Hard Can It Be? Hardness-Aware Multi-Objective Unlearning
Jiangwei Chen, Xinyuan Niu, Rachael Hwee Ling Sim +3
Machine unlearning aims to remove the influence of specific forget training data due to privacy, copyright or bias concerns while maintaining the model performance on the remaining…
Incentivizing Truthfulness and Collaborative Fairness in Bayesian Learning
Rachael Hwee Ling Sim, Jue Fan, Xiao Tian +3
Collaborative machine learning involves training high-quality models using datasets from a number of sources. To incentivize sources to share data, existing data valuation methods…
Is Data Shapley Not Better than Random in Data Selection? Ask NASH
Xiao Tian, Jue Fan, Rachael Hwee Ling Sim +3
Data selection studies the problem of identifying high-quality subsets of training data. While some existing works have considered selecting the subset of data with top- Data Sh…
INO-SGD: Addressing Utility Imbalance under Individualized Differential Privacy
Xiao Tian, Jue Fan, Rachael Hwee Ling Sim +1
Differential privacy (DP) is widely employed in machine learning to protect confidential or sensitive training data from being revealed. As data owners gain greater control over th…
WaterDrum: Watermarking for Data-centric Unlearning Metric
Xinyang Lu, Xinyuan Niu, Gregory Kang Ruey Lau +7
Large language model (LLM) unlearning is critical in real-world applications where it is necessary to efficiently remove the influence of private, copyrighted, or harmful data from…