4 papers · 1 filter
From Implicit to Explicit: Enhancing Self-Recognition in Large Language Models
Yinghan Zhou, Weifeng Zhu, Juan Wen +3
Large language models (LLMs) have been shown to possess a degree of self-recognition ability, which used to identify whether a given text was generated by themselves. Prior work ha…
BadApex: Backdoor Attack Based on Adaptive Optimization Mechanism of Black-box Large Language Models
Zhengxian Wu, Juan Wen, Wanli Peng +3
Previous insertion-based and paraphrase-based backdoors have achieved great success in attack efficacy, but they ignore the text quality and semantic consistency between poisoned a…
Kill two birds with one stone: generalized and robust AI-generated text detection via dynamic perturbations
Yinghan Zhou, Juan Wen, Wanli Peng +3
The growing popularity of large language models has raised concerns regarding the potential to misuse AI-generated text (AIGT). It becomes increasingly critical to establish an exc…
ImF: Implicit Fingerprint for Large Language Models
Jiaxuan Wu, Wanli Peng, Hang Fu +2
Training large language models (LLMs) is resource-intensive and expensive, making protecting intellectual property (IP) for LLMs crucial. Recently, embedding fingerprints into LLMs…