12 papers
SLIP: Soft Label Mechanism and Key-Extraction-Guided CoT-based Defense Against Instruction Backdoor in APIs
Zhengxian Wu, Juan Wen, Wanli Peng +3
Customized Large Language Model (LLM) agents face a critical security threat from black-box instruction backdoors, where malicious behaviors are covertly injected through hidden sy…
Protecting Creative Writing Copyright against AI Imitation via Implicit Watermarking
Ziwei Zhang, Juan Wen, Wanli Peng +3
Large language models (LLMs) enable powerful knowledge injection through approaches such as in-context learning and fine-tuning, but they also introduce new risks of unauthorized i…
BeDKD: Backdoor Defense Based on Directional Mapping Module and Adversarial Knowledge Distillation
Zhengxian Wu, Juan Wen, Wanli Peng +3
Although existing backdoor defenses have gained success in mitigating backdoor attacks, they still face substantial challenges. In particular, most of them rely on large amounts of…
From Implicit to Explicit: Enhancing Self-Recognition in Large Language Models
Yinghan Zhou, Weifeng Zhu, Juan Wen +3
Large language models (LLMs) have been shown to possess a degree of self-recognition ability, which used to identify whether a given text was generated by themselves. Prior work ha…
Inhibitory Attacks on Backdoor-based Fingerprinting for Large Language Models
Hang Fu, Wanli Peng, Yinghan Zhou +3
The widespread adoption of Large Language Model (LLM) in commercial and research settings has intensified the need for robust intellectual property protection. Backdoor-based LLM f…
ImF: Implicit Fingerprint for Large Language Models
Jiaxuan Wu, Wanli Peng, Hang Fu +2
Training large language models (LLMs) is resource-intensive and expensive, making protecting intellectual property (IP) for LLMs crucial. Recently, embedding fingerprints into LLMs…