68 citations · 74 across the 6 of their papers we have counts for
1 paper · 1 filter
Zhongyang Lin, Ziran Zhao, Feifei Zhai +1
Large language models remain vulnerable to jailbreak attacks that hide harmful intent behind seemingly ordinary requests such as role-play, translation, encoding, adversarial suffi…