241 citations · 838 across the 17 of their papers we have counts for
5 papers · 1 filter
Adversarial Demonstration Attacks on Large Language Models
Jiongxiao Wang, Zichen Liu, Keun Hee Park +5
With the emergence of more powerful large language models (LLMs), such as ChatGPT and GPT-4, in-context learning (ICL) has gained significant prominence in leveraging these models…
From Shortcuts to Triggers: Backdoor Defense with Denoised PoE
Qin Liu, Fei Wang, Chaowei Xiao +1
Language models are often at risk of diverse backdoor attacks, especially data poisoning. Thus, it is important to investigate defense solutions for addressing them. Existing backd…
Instructions as Backdoors: Backdoor Vulnerabilities of Instruction Tuning for Large Language Models
Jiashu Xu, Mingyu Derek Ma, Fei Wang +2
We investigate security concerns of the emergent instruction tuning paradigm, that models are trained on crowdsourced datasets with task instructions to achieve superior performanc…
Defending against Insertion-based Textual Backdoor Attacks via Attribution
Jiazhao Li, Zhuofeng Wu, Wei Ping +2
Textual backdoor attack, as a novel attack model, has been shown to be effective in adding a backdoor to the model during training. Defending against such backdoor attacks has beco…
Shall We Pretrain Autoregressive Language Models with Retrieval? A Comprehensive Study
Boxin Wang, Wei Ping, Peng Xu +9
Large decoder-only language models (LMs) can be largely improved in terms of perplexity by retrieval (e.g., RETRO), but its impact on text generation quality and downstream task ac…