7 papers
How Faithful Is Trajectory-Based Data Attribution? Error Sources, Remedies, and Practical Guidelines
Junwei Deng, Pingbang Hu, Suliang Jin +4
Trajectory-based data attribution methods estimate the influence of training samples on model predictions by unrolling the training trajectory. They are widely used in applications…
Dr. Post-Training: A Data Regularization Perspective on LLM Post-Training
Pingbang Hu, Xueshen Liu, Z. Morley Mao +1
Data selection methods address a critical challenge in LLM post-training: effectively leveraging scarce, high-fidelity target data alongside abundant but imperfectly aligned genera…
A Unified Theory of Random Projection for Influence Functions
Pingbang Hu, Yuzheng Hu, Jiaqi W. Ma +1
Influence functions and related data attribution scores take the form of , where is a curvature operator. In modern overparametrized models,…
A Reliable Cryptographic Framework for Empirical Machine Unlearning Evaluation
Yiwen Tu, Pingbang Hu, Jiaqi Ma
Machine unlearning updates machine learning models to remove information from specific training samples, complying with data protection regulations that allow individuals to reques…
GraSS: Scalable Data Attribution with Gradient Sparsification and Sparse Projection
Pingbang Hu, Joseph Melkonian, Weijing Tang +2
Gradient-based data attribution methods, such as influence functions, are critical for understanding the impact of individual training samples without requiring repeated model retr…
Adversarial Attacks on Data Attribution
Xinhe Wang, Pingbang Hu, Junwei Deng +1
Data attribution aims to quantify the contribution of individual training data points to the outputs of an AI model, which has been used to measure the value of training data and c…