35 citations · 36 across the 17 of their papers we have counts for
1 paper · 2 filters
Marco Rando, Samuel Vaiter
Large language models (LLMs) are known to be vulnerable to jailbreak attacks, which typically rely on carefully designed prompts containing explicit semantic structure. These attac…