1 citations · 2 across the 5 of their papers we have counts for
3 papers · 1 filter
PandaGuard: Systematic Evaluation of LLM Safety against Jailbreaking Attacks
Guobin Shen, Dongcheng Zhao, Linghao Feng +8
Large language models (LLMs) have achieved remarkable capabilities but remain vulnerable to adversarial prompts known as jailbreaks, which can bypass safety alignment and elicit ha…
Harnessing Task Overload for Scalable Jailbreak Attacks on Large Language Models
Yiting Dong, Guobin Shen, Dongcheng Zhao +2
Large Language Models (LLMs) remain vulnerable to jailbreak attacks that bypass their safety mechanisms. Existing attack methods are fixed or specifically tailored for certain mode…
Jailbreak Antidote: Runtime Safety-Utility Balance via Sparse Representation Adjustment in Large Language Models
Guobin Shen, Dongcheng Zhao, Yiting Dong +2
As large language models (LLMs) become integral to various applications, ensuring both their safety and utility is paramount. Jailbreak attacks, which manipulate LLMs into generati…