1 paper
Muyang Zheng, Yuanzhi Yao, Changting Lin +3
Despite efforts to align large language models (LLMs) with societal and moral values, these models remain susceptible to jailbreak attacks -- methods designed to elicit harmful res…