Showing cs.LGShow all
2 papers · 1 filter
cs.LG2025
Final-Model-Only Data Attribution with a Unifying View of Gradient-Based Methods
Dennis Wei, Inkit Padhi, Soumya Ghosh +3
Training data attribution (TDA) is concerned with understanding model behavior in terms of the training data. This paper draws attention to the common setting where one has access…
cs.LG2024
Split, Unlearn, Merge: Leveraging Data Attributes for More Effective Unlearning in LLMs
Swanand Ravindra Kadhe, Farhan Ahmed, Dennis Wei +2
Large language models (LLMs) have shown to pose social and ethical risks such as generating toxic language or facilitating malicious use of hazardous knowledge. Machine unlearning…