5 papers
Rethinking Backdoor Adversarial Unlearning through the Lens of Catastrophic Forgetting in Continual Learning
Zhenqian Zhu, Yamin Hu, Yujiang Liu +5
Existing studies reveal that current backdoor defenses exhibit limited robustness and often fail against specific types of attacks. More concerningly, prevailing safety tuning stra…
From Parameters to Feature Space: Task Arithmetic for Backdoor Mitigation in Model Merging
Zhenqian Zhu, Yamin Hu, Yiya Diao +3
Model merging (MM) has gained significant attention as a cost-effective approach to integrate multiple task-specific models into a unified model. However, recent work reveals that…
Acoda: Adversarial Code Obfuscation for Defending against LLM-based Analysis
Hongzhou Rao, Zikan Dong, Yanjie Zhao +2
With the widespread adoption of Large Language Models (LLMs) in software engineering (SE) tasks such as code understanding, debugging, and vulnerability detection, their powerful s…
Repurposing and Evaluating the (In)Feasibility of Dataset Poisoning enabled Watermarking for Contrastive Learning
Zhiyang Dai, Yansong Gao, Boyu Kuang +5
Contrastive learning (CL) reduces annotation cost via auto-derived supervisory signals. Since large-scale in-house CL datasets are infeasible, reliance on third-party or internet d…
As If We've Met Before: LLMs Exhibit Certainty in Recognizing Seen Files
Haodong Li, Jingqi Zhang, Xiao Cheng +3
The remarkable language ability of Large Language Models (LLMs) stems from extensive training on vast datasets, often including copyrighted material, which raises serious concerns…