1 paper · 1 filter
Mario RodrÃguez Béjar, Francisco J. Cortés-Delgado, S. Braghin +1
Large language models (LLMs) remain vulnerable to jailbreak attacks that bypass safety alignment and elicit harmful responses. A growing body of work shows that contextual priming,…