1 paper
Lixing Lin, Juli You, Yue Li +4
Large language model (LLM) safety classifiers such as Llama Guard are effective at detecting overtly harmful prompts but remain vulnerable to adversarial jailbreak attacks that dis…