14 citations · 23 across the 15 of their papers we have counts for
Showing 2024 · cs.CRShow all
2 papers · 2 filters
cs.CR2024
Harnessing Task Overload for Scalable Jailbreak Attacks on Large Language Models
Yiting Dong, Guobin Shen, Dongcheng Zhao +2
Large Language Models (LLMs) remain vulnerable to jailbreak attacks that bypass their safety mechanisms. Existing attack methods are fixed or specifically tailored for certain mode…
cs.CR2024★ 4 cited
Jailbreak Antidote: Runtime Safety-Utility Balance via Sparse Representation Adjustment in Large Language Models
Guobin Shen, Dongcheng Zhao, Yiting Dong +2
As large language models (LLMs) become integral to various applications, ensuring both their safety and utility is paramount. Jailbreak attacks, which manipulate LLMs into generati…