1 paper
Zhongyang Lin, Ziran Zhao, Feifei Zhai +1
Large language models remain vulnerable to jailbreak attacks that hide harmful intent behind seemingly ordinary requests such as role-play, translation, encoding, adversarial suffi…