1 citations · 1 across the 1 of their papers we have counts for
1 paper
Guobin Shen, Dongcheng Zhao, Linghao Feng +8
Large language models (LLMs) have achieved remarkable capabilities but remain vulnerable to adversarial prompts known as jailbreaks, which can bypass safety alignment and elicit ha…