collaborators

10 papers

cs.CR2026

SLIP: Soft Label Mechanism and Key-Extraction-Guided CoT-based Defense Against Instruction Backdoor in APIs

Zhengxian Wu, Juan Wen, Wanli Peng +3

Customized Large Language Model (LLM) agents face a critical security threat from black-box instruction backdoors, where malicious behaviors are covertly injected through hidden sy…

cs.CR2026

Protecting Creative Writing Copyright against AI Imitation via Implicit Watermarking

Ziwei Zhang, Juan Wen, Wanli Peng +3

Large language models (LLMs) enable powerful knowledge injection through approaches such as in-context learning and fine-tuning, but they also introduce new risks of unauthorized i…

cs.CR2026

BeDKD: Backdoor Defense Based on Directional Mapping Module and Adversarial Knowledge Distillation

Zhengxian Wu, Juan Wen, Wanli Peng +3

Although existing backdoor defenses have gained success in mitigating backdoor attacks, they still face substantial challenges. In particular, most of them rely on large amounts of…

cs.CL2026

From Implicit to Explicit: Enhancing Self-Recognition in Large Language Models

Yinghan Zhou, Weifeng Zhu, Juan Wen +3

Large language models (LLMs) have been shown to possess a degree of self-recognition ability, which used to identify whether a given text was generated by themselves. Prior work ha…

cs.CR2026

Inhibitory Attacks on Backdoor-based Fingerprinting for Large Language Models

Hang Fu, Wanli Peng, Yinghan Zhou +3

The widespread adoption of Large Language Model (LLM) in commercial and research settings has intensified the need for robust intellectual property protection. Backdoor-based LLM f…

cs.CR2025

Self-Disguise Attack: Induce the LLM to disguise itself for AIGT detection evasion

Yinghan Zhou, Juan Wen, Wanli Peng +3

AI-generated text (AIGT) detection evasion aims to reduce the detection probability of AIGT, helping to identify weaknesses in detectors and enhance their effectiveness and reliabi…