1 paper
Ruben Belo, Marta Guimaraes, Claudia Soares
Large Language Models are susceptible to jailbreak attacks that bypass built-in safety guardrails (e.g., by tricking the model with adversarial prompts). We propose Concept Alignme…