1 paper
Shiyu Xiang, Ansen Zhang, Yanfei Cao +2
Although Aligned Large Language Models (LLMs) are trained to refuse harmful requests, they remain vulnerable to jailbreak attacks. Unfortunately, existing methods often focus on su…