9 citations · 31 across the 42 of their papers we have counts for
4 papers · 1 filter
Bidirectional Intention Inference Enhances LLMs' Defense Against Multi-Turn Jailbreak Attacks
Haibo Tong, Dongcheng Zhao, Guobin Shen +4
The remarkable capabilities of Large Language Models (LLMs) have raised significant safety concerns, particularly regarding "jailbreak" attacks that exploit adversarial prompts to…
PandaGuard: Systematic Evaluation of LLM Safety against Jailbreaking Attacks
Guobin Shen, Dongcheng Zhao, Linghao Feng +8
Large language models (LLMs) have achieved remarkable capabilities but remain vulnerable to adversarial prompts known as jailbreaks, which can bypass safety alignment and elicit ha…
Harnessing Task Overload for Scalable Jailbreak Attacks on Large Language Models
Yiting Dong, Guobin Shen, Dongcheng Zhao +2
Large Language Models (LLMs) remain vulnerable to jailbreak attacks that bypass their safety mechanisms. Existing attack methods are fixed or specifically tailored for certain mode…
Jailbreak Antidote: Runtime Safety-Utility Balance via Sparse Representation Adjustment in Large Language Models
Guobin Shen, Dongcheng Zhao, Yiting Dong +2
As large language models (LLMs) become integral to various applications, ensuring both their safety and utility is paramount. Jailbreak attacks, which manipulate LLMs into generati…