1 citations · 2 across the 11 of their papers we have counts for
Showing 2024 · cs.CRShow all
3 papers · 2 filters
cs.CR2024
ASPIRER: Bypassing System Prompts With Permutation-based Backdoors in LLMs
Lu Yan, Siyuan Cheng, Xuan Chen +4
Large Language Models (LLMs) have become integral to many applications, with system prompts serving as a key mechanism to regulate model behavior and ensure ethical outputs. In thi…
cs.CR2024
RL-JACK: Reinforcement Learning-powered Black-box Jailbreaking Attack against LLMs
Xuan Chen, Yuzhou Nie, Lu Yan +3
Modern large language model (LLM) developers typically conduct a safety alignment to prevent an LLM from generating unethical or harmful content. Recent studies have discovered tha…
cs.CR2024★ 1 cited
When LLM Meets DRL: Advancing Jailbreaking Efficiency via DRL-guided Search
Xuan Chen, Yuzhou Nie, Wenbo Guo +1
Recent studies developed jailbreaking attacks, which construct jailbreaking prompts to fool LLMs into responding to harmful questions. Early-stage jailbreaking attacks require acce…