1 paper · 1 filter
Haibo Jin, Andy Zhou, Joe D. Menke +1
Large Language Models (LLMs) are typically harmless but remain vulnerable to carefully crafted prompts known as ``jailbreaks'', which can bypass protective measures and induce harm…