5 papers
Who Wrote the Book? Detecting and Attributing LLM Ghostwriters
Anudeex Shetty, Qiongkai Xu, Olga Ohrimenko +1
In this paper, we introduce GhostWriteBench, a dataset for LLM authorship attribution. It comprises long-form texts (50K+ words per book) generated by frontier LLMs, and is designe…
Defending against Backdoor Attacks via Module Switching
Weijun Li, Ansh Arora, Xuanli He +2
Backdoor attacks pose a serious threat to deep neural networks (DNNs), allowing adversaries to implant triggers for hidden behaviors in inference. Defending against such vulnerabil…
Cut the Deadwood Out: Backdoor Purification via Guided Module Substitution
Yao Tong, Weijun Li, Xuanli He +2
Model NLP models are commonly trained (or fine-tuned) on datasets from untrusted platforms like HuggingFace, posing significant risks of data poisoning attacks. A practical yet und…
GRADA: Graph-based Reranking against Adversarial Documents Attack
Jingjie Zheng, Aryo Pradipta Gema, Giwon Hong +4
Retrieval Augmented Generation (RAG) frameworks improve the accuracy of large language models (LLMs) by integrating external knowledge from retrieved documents, thereby overcoming…
TuBA: Cross-Lingual Transferability of Backdoor Attacks in LLMs with Instruction Tuning
Xuanli He, Jun Wang, Qiongkai Xu +4
The implications of backdoor attacks on English-centric large language models (LLMs) have been widely examined - such attacks can be achieved by embedding malicious behaviors durin…