27 citations · 27 across the 2 of their papers we have counts for
1 paper · 1 filter
Peng Ding, Jun Kuang, Wen Sun +5
Large language models (LLMs) remain vulnerable to jailbreaking attacks despite their impressive capabilities. Investigating these weaknesses is crucial for robust safety mechanisms…