activity
20242026
collaborators
Showing cs.LGShow all

6 papers · 1 filter

cs.LG2026

Sequential Data Poisoning in LLM Post-Training

Jack Sanderson, Yihan Wang, Xiaoqian Lu +2

LLM post-training proceeds through multiple stages, e.g., supervised fine-tuning (SFT) followed by reinforcement learning from human feedback (RLHF) or direct preference optimizati…

cs.LG2026

Are Targeted Data Poisoning Attacks as Effective as We Think?

William Xu, Chenyu Zhang, Yihan Wang +5

Targeted data poisoning attacks manipulate model predictions on specific test samples by injecting malicious data into training. Yet existing evaluations report average attack succ…

cs.LG2026

Machine Unlearning Fails to Remove Data Poisoning Attacks

Martin Pawelczyk, Jimmy Z. Di, Yiwei Lu +3

We revisit the efficacy of several practical methods for approximate machine unlearning developed for large-scale deep learning. In addition to complying with data deletion request…

cs.LG2025

MUC: Machine Unlearning for Contrastive Learning with Black-box Evaluation

Yihan Wang, Yiwei Lu, Guojun Zhang +4

Machine unlearning offers effective solutions for revoking the influence of specific training data on pre-trained model parameters. While existing approaches address unlearning for…

cs.LG2025

BridgePure: Limited Protection Leakage Can Break Black-Box Data Protection

Yihan Wang, Yiwei Lu, Xiao-Shan Gao +2

Availability attacks, or unlearnable examples, are defensive techniques that allow data owners to modify their datasets in ways that prevent unauthorized machine learning models fr…

cs.LG2024

Disguised Copyright Infringement of Latent Diffusion Models

Yiwei Lu, Matthew Y. R. Yang, Zuoqiu Liu +2

Copyright infringement may occur when a generative model produces samples substantially similar to some copyrighted data that it had access to during the training phase. The notion…