2 papers
cs.LG2025
IF-GUIDE: Influence Function-Guided Detoxification of LLMs
Zachary Coalson, Juhan Bae, Nicholas Carlini +1
We study how training data contributes to the emergence of toxic behaviors in large language models. Most prior work on reducing model toxicity adopts reactive approaches, such as…
cs.CR2025
Certified Robustness to Clean-Label Poisoning Using Diffusion Denoising
Sanghyun Hong, Nicholas Carlini, Alexey Kurakin
We present a certified defense to clean-label poisoning attacks under -norm. These attacks work by injecting a small number of poisoning samples (e.g., 1%) that contain bou…