18 citations · 90 across the 17 of their papers we have counts for
Showing cs.CRShow all
3 papers · 1 filter
cs.CR2024★ 1 cited
FATH: Authentication-based Test-time Defense against Indirect Prompt Injection Attacks
Jiongxiao Wang, Fangzhou Wu, Wendi Li +5
Large language models (LLMs) have been widely deployed as the backbone with additional tools and text information for real-world applications. However, integrating external informa…
cs.CR2024★ 3 cited
Mitigating Fine-tuning based Jailbreak Attack with Backdoor Enhanced Safety Alignment
Jiongxiao Wang, Jiazhao Li, Yiquan Li +7
Despite the general capabilities of Large Language Models (LLM), these models still request fine-tuning or adaptation with customized data when meeting specific business demands. H…
cs.CR2023★ 14 cited
On the Exploitability of Instruction Tuning
Manli Shu, Jiongxiao Wang, Chen Zhu +3
Instruction tuning is an effective technique to align large language models (LLMs) with human intents. In this work, we investigate how an adversary can exploit instruction tuning…