33 citations · 33 across the 1 of their papers we have counts for
2 papers
cs.CR2025
Why LLM Safety Guardrails Collapse After Fine-tuning: A Similarity Analysis Between Alignment and Fine-tuning Datasets
Lei Hsiung, Tianyu Pang, Yung-Chen Tang +4
Recent advancements in large language models (LLMs) have underscored their vulnerability to safety alignment jailbreaks, particularly when subjected to downstream fine-tuning. Howe…
cs.CR2022★ 33 cited
Neurotoxin: Durable Backdoors in Federated Learning
Zhengming Zhang, Ashwinee Panda, Linyue Song +5
Due to their decentralized nature, federated learning (FL) systems have an inherent vulnerability during their training to adversarial backdoor attacks. In this type of attack, the…