1 paper · 1 filter
Aldan Creo, Raul Castro Fernandez, Manuel Cebrian
As large language models (LLMs) become increasingly deployed, understanding the complexity and evolution of jailbreaking strategies is critical for AI safety. We present a mass-sca…