1 paper
Zheng Lin, Zhenxing Niu, Haoxuan Ji +2
This paper proposes a jailbreaking prompt detection method for large language models (LLMs) to defend against jailbreak attacks. Although recent LLMs are equipped with built-in saf…