most citedFrom Implicit to Explicit: Enhancing Self-Recognition in Large Language Models

1 citations · 1 across the 7 of their papers we have counts for

collaborators

8 papers

cs.CR2025

Self-Disguise Attack: Induce the LLM to disguise itself for AIGT detection evasion

Yinghan Zhou, Juan Wen, Wanli Peng +3

AI-generated text (AIGT) detection evasion aims to reduce the detection probability of AIGT, helping to identify weaknesses in detectors and enhance their effectiveness and reliabi…

cs.CL20251 cited

From Implicit to Explicit: Enhancing Self-Recognition in Large Language Models

Yinghan Zhou, Weifeng Zhu, Juan Wen +3

Large language models (LLMs) have been shown to possess a degree of self-recognition ability, which used to identify whether a given text was generated by themselves. Prior work ha…

cs.CR2025

SLIP: Soft Label Mechanism and Key-Extraction-Guided CoT-based Defense Against Instruction Backdoor in APIs

Zhengxian Wu, Juan Wen, Wanli Peng +3

Customized Large Language Model (LLM) agents face a critical security threat from black-box instruction backdoors, where malicious behaviors are covertly injected through hidden sy…

cs.CR2025

BeDKD: Backdoor Defense Based on Directional Mapping Module and Adversarial Knowledge Distillation

Zhengxian Wu, Juan Wen, Wanli Peng +3

Although existing backdoor defenses have gained success in mitigating backdoor attacks, they still face substantial challenges. In particular, most of them rely on large amounts of…

cs.CR2025

GTSD: Generative Text Steganography Based on Diffusion Model

Zhengxian Wu, Juan Wen, Yiming Xue +2

With the rapid development of deep learning, existing generative text steganography methods based on autoregressive models have achieved success. However, these autoregressive steg…

cs.CL2025

BadApex: Backdoor Attack Based on Adaptive Optimization Mechanism of Black-box Large Language Models

Zhengxian Wu, Juan Wen, Wanli Peng +3

Previous insertion-based and paraphrase-based backdoors have achieved great success in attack efficacy, but they ignore the text quality and semantic consistency between poisoned a…