1 paper
Rachneet Sachdeva, Rima Hazra, Iryna Gurevych
Large language models, despite extensive alignment with human values and ethical principles, remain vulnerable to sophisticated jailbreak attacks that exploit their reasoning abili…