13 papers
REED: Post-Training Representation Editing for Cross-Domain Linguistic Steganalysis
Ruohan Lei, Jianxin Gao, Wanli Peng +1
In real-world scenarios of linguistic steganalysis, tested texts usually come from unseen domains with different vocabularies, topics, writing styles, and steganographic generation…
Text Steganography with Dynamic Codebook and Multimodal Large Language Model
Jianxin Gao, Ruohan Lei, Wanli Peng
With the popularity of the large language models (LLMs), text steganography has achieved remarkable performance. However, existing methods still have some issues: (1) For the white…
SLIP: Soft Label Mechanism and Key-Extraction-Guided CoT-based Defense Against Instruction Backdoor in APIs
Zhengxian Wu, Juan Wen, Wanli Peng +3
Customized Large Language Model (LLM) agents face a critical security threat from black-box instruction backdoors, where malicious behaviors are covertly injected through hidden sy…
Protecting Creative Writing Copyright against AI Imitation via Implicit Watermarking
Ziwei Zhang, Juan Wen, Wanli Peng +3
Large language models (LLMs) enable powerful knowledge injection through approaches such as in-context learning and fine-tuning, but they also introduce new risks of unauthorized i…
BeDKD: Backdoor Defense Based on Directional Mapping Module and Adversarial Knowledge Distillation
Zhengxian Wu, Juan Wen, Wanli Peng +3
Although existing backdoor defenses have gained success in mitigating backdoor attacks, they still face substantial challenges. In particular, most of them rely on large amounts of…
From Implicit to Explicit: Enhancing Self-Recognition in Large Language Models
Yinghan Zhou, Weifeng Zhu, Juan Wen +3
Large language models (LLMs) have been shown to possess a degree of self-recognition ability, which used to identify whether a given text was generated by themselves. Prior work ha…