1 paper
Feiyue Xu, Hongsheng Hu, Chaoxiang He +9
Large Language Models (LLMs) have achieved remarkable success but remain highly susceptible to jailbreak attacks, in which adversarial prompts coerce models into generating harmful…