1 paper · 1 filter
Nathalie Kirch, Constantin Weisser, Severin Field +2
Jailbreaks have been a central focus of research regarding the safety and reliability of large language models (LLMs), yet the mechanisms underlying these attacks remain poorly und…