1 paper · 1 filter
Hanqing Liu, Lifeng Zhou, Huanqian Yan
Large language models have drawn significant attention to the challenge of safe alignment, especially regarding jailbreak attacks that circumvent security measures to produce harmf…