1 citations · 1 across the 1 of their papers we have counts for
3 papers
cs.LG2025
IF-GUIDE: Influence Function-Guided Detoxification of LLMs
Zachary Coalson, Juhan Bae, Nicholas Carlini +1
We study how training data contributes to the emergence of toxic behaviors in large language models. Most prior work on reducing model toxicity adopts reactive approaches, such as…
cs.CR2024★ 1 cited
Privacy Backdoors: Enhancing Membership Inference through Poisoning Pre-trained Models
Yuxin Wen, Leo Marchyok, Sanghyun Hong +3
It is commonplace to produce application-specific models by fine-tuning large pre-trained models using a small bespoke dataset. The widespread availability of foundation model chec…
cs.CR2024
Certified Robustness to Clean-Label Poisoning Using Diffusion Denoising
Sanghyun Hong, Nicholas Carlini, Alexey Kurakin
We present a certified defense to clean-label poisoning attacks under -norm. These attacks work by injecting a small number of poisoning samples (e.g., 1%) that contain bou…