1 citations · 1 across the 2 of their papers we have counts for
1 paper · 1 filter
Yanshu Wang, Shuaishuai Yang, Jingjing He +1
Large Language Models (LLMs) face increasing threats from jailbreak attacks that bypass safety alignment. While prompt-based defenses such as Role-Oriented Prompts (RoP) and Task-O…